Inference Engineering — 13 · Month Roadmap¶
LLM serving, GPU kernels, quantization, and distributed inference.
00 · Command
- 00_COMMAND — The Master Layer
- 01 — Month-by-Month Plan (M1 → M13)
- 02 — Sprint Calendar (S1 → S26)
- 03 — North-Star Artifacts (The Eight Rungs)
- 04 — Weekly Rhythm (The Canonical Week)
- 05 — KPI Dashboard (Weekly, Monthly, Quarterly)
- 06 — Zoho Alignment (Turning Your Day Job Into Your Portfolio)
- 07 — The 13-Month Pitch (Who You Are At M13)
01 · Foundations
02 · Transformers
- Phase 1 — Transformers Down to the Metal
- 01 — The Karpathy Path (Zero to Hero → nanoGPT → llm.c → nanochat)
- 02 — From-Scratch Inference (Round 2 Build)
- 03 — The Arithmetic of Transformers (The Killer Notebook)
- 04 — Stanford CS336: Language Modeling from Scratch (Your Backbone)
- 05 — The Paper Canon for Phase 1
- 06 — Phase 1 Projects (What You Ship)
03 · GPU Kernels
- Phase 2 — GPU Architecture & Kernel Programming
- The GPU Hardware Model
- The PMPP Path
- GPU MODE — the community
- The SGEMM Ladder
- Triton — the pragmatic kernel language for inference
- CUTLASS + CuTe — the industrial kernel language
- Profiling Mastery — Nsight + torch.profiler
- AMD ROCm/HIP + Apple MLX — awareness-level
- Phase 2 — Projects (with acceptance criteria)
04 · Attention
- Phase 3 — Attention & Fusion Kernels
- Online Softmax — the mathematical heart of FlashAttention
- The FlashAttention lineage — FA1 / FA2 / FA3 / FA4
- FlashDecoding + FlashInfer — the decode-time attention story
- Writing FlashAttention 2 in Triton — a step-by-step build
- Fusion Thinking — the second-most-valuable inference skill
- Numerics discipline for low-precision inference
- Phase 3 Projects — concrete artifacts with acceptance criteria
05 · Engines
- Phase 4 — Inference Engine Internals
- 01 — The Metrics Language of Serving
- 02 — Orca & Continuous Batching
- 03 — PagedAttention & the vLLM V1 Architecture
- 04 — Chunked Prefill & Stall-Free Scheduling (Sarathi-Serve)
- 05 — RadixAttention & SGLang (YOUR PAPER)
- 06 — Speculative Decoding: The Full Lineage
- 07 — Structured / Guided Decoding
- 08 — CUDA Graphs in vLLM
- 09 — Mini Inference Engine (The Capstone)
- 10 — vLLM V1 Source Map (Code Is Truth)
- 11 — SGLang Source Map
- 12 — llama.cpp / GGUF / Local Inference Mastery
- 13 — The Local Inference Ecosystem
- 14 — Phase 4 Projects (Acceptance Criteria)
06 · Quantization
- Phase 5 — Quantization & Model Compression
- 01 — Quantization Theory: Affine Maps, Granularity, and the Dequant Math
- 02 — The Outlier Problem: Why Naive INT8 Falls Off a Cliff on LLMs
- 03 — The Asymmetry Rule: W4A16 vs W8A8/FP8, With Worked Roofline Math
- 04 — GPTQ: Hessian-Weighted Rounding, the OBS Lineage, and Why It Still Ships
- 05 — AWQ: Activation-Aware Weight Scaling
- 06 — SmoothQuant: Migrating Difficulty from Activations to Weights
- 07 — The FP8 Family: E4M3, E5M2, and Per-Block Scaling
- 08 — FP4, MXFP4, NVFP4: The Blackwell Frontier
- 09 — KV Cache Quantization
- 10 — GGUF k-quants and i-quants: The Bit-Level Reality
- 11 — Marlin, Machete, and the Quantized-GEMM Kernel Reality
- 12 — Evaluation Discipline: No Eval = Vandalism
- 13 — THE Bake-Off Project: One 8B Model, Four Formats, Two Tables
- 14 — Phase 5 Projects Portfolio
07 · Distributed
- Phase 6 — Distributed Inference & Training at Scale
- 01 — The 5D Parallelism Taxonomy
- 02 — Tensor Parallelism
- 03 — Pipeline Parallelism
- 04 — ZeRO and FSDP
- 05 — Expert Parallelism (MoE)
- 06 — Sequence / Context Parallelism (Ring Attention)
- 07 — NCCL & The Network Fabric
- 08 — Distributed Inference in vLLM & SGLang
- 09 — Disaggregated Prefill/Decode Serving
- 10 — KV-Cache-Aware Routing: Why Scheduling Beats Kernel Optimization at Fleet Scale
- 11 — The HuggingFace Ultra-Scale Playbook: A Guided Tour
- 12 — “How to Scale Your Model” (Google DeepMind / JAX Scaling Book)
- 13 — Pretraining a Small Model End-to-End (124M → 1B)
- 14 — modded-nanogpt: the Speedrun Community
- 15 — The Fine-Tuning Stack: LoRA, QLoRA, DPO, GRPO
- 16 — Rollout Infrastructure: The Training Loop That Contains an Inference Engine
- 17 — The DeepSeek-V3 & R1 Technical Reports: A Guided Reading
- 18 — Phase 6 Projects: The Distributed-Scale Portfolio
08 · Production
- Phase 7 — Production & Enterprise-Grade Inference
- 01 — Engines as Products
- 02 — The OpenAI-Compatible API as Integration Contract
- 03 — Kubernetes for GPUs
- 04 — Autoscaling LLM Workloads
- 05 — Observability for LLM Inference
- 06 — Reliability: OOMs, Preemption, Backpressure, Multi-Tenancy, Abuse
- 07 — On-Prem & Enterprise: The Zoho Unfair Advantage Manifesto
- 08 — Sizing Exercises: The Napkin Math That Wins Deals
- 09 — Cost Modeling: $/1M Tokens, Tiering, and Fine-Tune-Small vs Prompt-Big
- 10 — GPU Procurement Literacy (2026)
- 11 — Security & Model Supply Chain
- 12 — Reference Architecture Capstone: On-Prem 8×H100 Enterprise Deployment
- 13 — Zoho Leverage Plan
- 14 — Phase 7 Projects (with acceptance criteria)
09 · Papers
10 · Communities
11 · Hardware
12 · Portfolio
- Portfolio Ladder: Signal-per-Unit-Effort
- Rung 1 — CPU Tiled Matmul Writeup
- Rung 2 — CUDA SGEMM Ladder
- Rung 3 — Triton FlashAttention-2
- Rung 4 — The Quantization Bake-Off
- Rung 5 — The Mini Inference Engine
- Rung 6 — The First Merged PR
- Rung 7 — The Enterprise Reference Architecture
- Rung 8 — The Sustained Contribution Area
13 · Discipline
99 · Pre Mortem
- 99 — Pre-Mortem: The Honest Version of the Plan
- 01 — The Industry Moves Faster Than the Plan
- 02 — Zoho Starves the Roadmap of Time
- 03 — The Valley of Despair (Months 6–8)
- 04 — Shiny Object Syndrome
- 05 — Open-Source Community & Political Risk
- 06 — Hardware Access Chokes At the Wrong Time
- Pre-Mortem 07 — Geography, Visa, and the Offer-Timing Risk
- Pre-Mortem 08 — The Unspeakable File
- Pre-Mortem 09 — Summary & Reset Protocol