Contents Menu Expand Light mode Dark mode Auto light/dark, in light mode Auto light/dark, in dark mode Skip to content
Inference Engineering — Project Fortress
Inference Engineering — Project Fortress

00 · Command

  • 00_COMMAND — The Master Layer
  • 01 — Month-by-Month Plan (M1 → M13)
  • 02 — Sprint Calendar (S1 → S26)
  • 03 — North-Star Artifacts (The Eight Rungs)
  • 04 — Weekly Rhythm (The Canonical Week)
  • 05 — KPI Dashboard (Weekly, Monthly, Quarterly)
  • 06 — Zoho Alignment (Turning Your Day Job Into Your Portfolio)
  • 07 — The 13-Month Pitch (Who You Are At M13)

01 · Foundations

  • Phase 0 — Systems & Math Foundations
  • 01 — Computer Systems: The Non-Negotiable Substrate
  • 02 — C++ for CUDA: The Pragmatic Subset
  • 03 — Python Performance Literacy
  • 04 — The Math (and Numerics) You Actually Need
  • 05 — Phase 0 Projects

02 · Transformers

  • Phase 1 — Transformers Down to the Metal
  • 01 — The Karpathy Path (Zero to Hero → nanoGPT → llm.c → nanochat)
  • 02 — From-Scratch Inference (Round 2 Build)
  • 03 — The Arithmetic of Transformers (The Killer Notebook)
  • 04 — Stanford CS336: Language Modeling from Scratch (Your Backbone)
  • 05 — The Paper Canon for Phase 1
  • 06 — Phase 1 Projects (What You Ship)

03 · GPU Kernels

  • Phase 2 — GPU Architecture & Kernel Programming
  • The GPU Hardware Model
  • The PMPP Path
  • GPU MODE — the community
  • The SGEMM Ladder
  • Triton — the pragmatic kernel language for inference
  • CUTLASS + CuTe — the industrial kernel language
  • Profiling Mastery — Nsight + torch.profiler
  • AMD ROCm/HIP + Apple MLX — awareness-level
  • Phase 2 — Projects (with acceptance criteria)

04 · Attention

  • Phase 3 — Attention & Fusion Kernels
  • Online Softmax — the mathematical heart of FlashAttention
  • The FlashAttention lineage — FA1 / FA2 / FA3 / FA4
  • FlashDecoding + FlashInfer — the decode-time attention story
  • Writing FlashAttention 2 in Triton — a step-by-step build
  • Fusion Thinking — the second-most-valuable inference skill
  • Numerics discipline for low-precision inference
  • Phase 3 Projects — concrete artifacts with acceptance criteria

05 · Engines

  • Phase 4 — Inference Engine Internals
  • 01 — The Metrics Language of Serving
  • 02 — Orca & Continuous Batching
  • 03 — PagedAttention & the vLLM V1 Architecture
  • 04 — Chunked Prefill & Stall-Free Scheduling (Sarathi-Serve)
  • 05 — RadixAttention & SGLang (YOUR PAPER)
  • 06 — Speculative Decoding: The Full Lineage
  • 07 — Structured / Guided Decoding
  • 08 — CUDA Graphs in vLLM
  • 09 — Mini Inference Engine (The Capstone)
  • 10 — vLLM V1 Source Map (Code Is Truth)
  • 11 — SGLang Source Map
  • 12 — llama.cpp / GGUF / Local Inference Mastery
  • 13 — The Local Inference Ecosystem
  • 14 — Phase 4 Projects (Acceptance Criteria)

06 · Quantization

  • Phase 5 — Quantization & Model Compression
  • 01 — Quantization Theory: Affine Maps, Granularity, and the Dequant Math
  • 02 — The Outlier Problem: Why Naive INT8 Falls Off a Cliff on LLMs
  • 03 — The Asymmetry Rule: W4A16 vs W8A8/FP8, With Worked Roofline Math
  • 04 — GPTQ: Hessian-Weighted Rounding, the OBS Lineage, and Why It Still Ships
  • 05 — AWQ: Activation-Aware Weight Scaling
  • 06 — SmoothQuant: Migrating Difficulty from Activations to Weights
  • 07 — The FP8 Family: E4M3, E5M2, and Per-Block Scaling
  • 08 — FP4, MXFP4, NVFP4: The Blackwell Frontier
  • 09 — KV Cache Quantization
  • 10 — GGUF k-quants and i-quants: The Bit-Level Reality
  • 11 — Marlin, Machete, and the Quantized-GEMM Kernel Reality
  • 12 — Evaluation Discipline: No Eval = Vandalism
  • 13 — THE Bake-Off Project: One 8B Model, Four Formats, Two Tables
  • 14 — Phase 5 Projects Portfolio

07 · Distributed

  • Phase 6 — Distributed Inference & Training at Scale
  • 01 — The 5D Parallelism Taxonomy
  • 02 — Tensor Parallelism
  • 03 — Pipeline Parallelism
  • 04 — ZeRO and FSDP
  • 05 — Expert Parallelism (MoE)
  • 06 — Sequence / Context Parallelism (Ring Attention)
  • 07 — NCCL & The Network Fabric
  • 08 — Distributed Inference in vLLM & SGLang
  • 09 — Disaggregated Prefill/Decode Serving
  • 10 — KV-Cache-Aware Routing: Why Scheduling Beats Kernel Optimization at Fleet Scale
  • 11 — The HuggingFace Ultra-Scale Playbook: A Guided Tour
  • 12 — “How to Scale Your Model” (Google DeepMind / JAX Scaling Book)
  • 13 — Pretraining a Small Model End-to-End (124M → 1B)
  • 14 — modded-nanogpt: the Speedrun Community
  • 15 — The Fine-Tuning Stack: LoRA, QLoRA, DPO, GRPO
  • 16 — Rollout Infrastructure: The Training Loop That Contains an Inference Engine
  • 17 — The DeepSeek-V3 & R1 Technical Reports: A Guided Reading
  • 18 — Phase 6 Projects: The Distributed-Scale Portfolio

08 · Production

  • Phase 7 — Production & Enterprise-Grade Inference
  • 01 — Engines as Products
  • 02 — The OpenAI-Compatible API as Integration Contract
  • 03 — Kubernetes for GPUs
  • 04 — Autoscaling LLM Workloads
  • 05 — Observability for LLM Inference
  • 06 — Reliability: OOMs, Preemption, Backpressure, Multi-Tenancy, Abuse
  • 07 — On-Prem & Enterprise: The Zoho Unfair Advantage Manifesto
  • 08 — Sizing Exercises: The Napkin Math That Wins Deals
  • 09 — Cost Modeling: $/1M Tokens, Tiering, and Fine-Tune-Small vs Prompt-Big
  • 10 — GPU Procurement Literacy (2026)
  • 11 — Security & Model Supply Chain
  • 12 — Reference Architecture Capstone: On-Prem 8×H100 Enterprise Deployment
  • 13 — Zoho Leverage Plan
  • 14 — Phase 7 Projects (with acceptance criteria)

09 · Papers

  • 09 — The Paper Canon
  • 01 — Foundations Papers
  • 02 — Kernels Papers
  • 03 — Engines Papers
  • 04 — Quantization Papers
  • 05 — Distributed Papers (Training + Serving at Scale)
  • 06 — Paper Reading Schedule (Month-by-Month, 13-Month Timeline)

10 · Communities

  • Community & Information Diet
  • GPU MODE
  • r/LocalLLaMA
  • Open Source Inference Engines: vLLM / SGLang / FlashInfer / llama.cpp
  • The Blog Canon
  • Twitter/X Follow List
  • YouTube Channels
  • Conferences: How to Consume Without Attending

11 · Hardware

  • Hardware Strategy for the 13-Month Timeline
  • Own GPU Choice: 3090 vs 4090 vs 5090 (2026 reality)
  • Rent GPU Strategy (2026 pricing reality)
  • Apple Silicon Lab (MLX + Metal)
  • Budget Scenarios: $50, $150, $500 per month
  • 2026 GPU Landscape (H100 / H200 / B200 / B300 / GB200 / MI300X / MI325X)

12 · Portfolio

  • Portfolio Ladder: Signal-per-Unit-Effort
  • Rung 1 — CPU Tiled Matmul Writeup
  • Rung 2 — CUDA SGEMM Ladder
  • Rung 3 — Triton FlashAttention-2
  • Rung 4 — The Quantization Bake-Off
  • Rung 5 — The Mini Inference Engine
  • Rung 6 — The First Merged PR
  • Rung 7 — The Enterprise Reference Architecture
  • Rung 8 — The Sustained Contribution Area

13 · Discipline

  • 13 — The Discipline Layer
  • 01 — Sprint Cadence
  • 02 — The Lab Notebook
  • 03 — Benchmark Hygiene as Identity
  • 04 — Reading Code Daily
  • 05 — Teach To Learn
  • 06 — Failure Modes
  • 07 — Motivation & Sustainment
  • 08 — Health & Burnout Prevention
  • 09 — study Conversion

99 · Pre Mortem

  • 99 — Pre-Mortem: The Honest Version of the Plan
  • 01 — The Industry Moves Faster Than the Plan
  • 02 — Zoho Starves the Roadmap of Time
  • 03 — The Valley of Despair (Months 6–8)
  • 04 — Shiny Object Syndrome
  • 05 — Open-Source Community & Political Risk
  • 06 — Hardware Access Chokes At the Wrong Time
  • Pre-Mortem 07 — Geography, Visa, and the Offer-Timing Risk
  • Pre-Mortem 08 — The Unspeakable File
  • Pre-Mortem 09 — Summary & Reset Protocol
Back to top
Copyright ©
Made with Furo