Twitter/X Follow List

30 handles. Curated. No filler.


The philosophy

Do NOT scroll Twitter’s main timeline. It will destroy your attention span and radicalize you politically. Instead:

  1. Create a private List called “Inference Eng” and add the 30 handles below.

  2. Access ONLY via that list URL. Bookmark it. Delete the main app.

  3. 15 min × 5 weekdays = 75 min/week. Scan for: paper releases, kernel tricks, hiring signals, engine features shipping.

  4. Never post political takes. Never quote-tweet drama. Never subtweet. This is a professional channel.

X handles change; verify by the person’s known org / pinned tweet before following. Handles below are known-good as of 2025–2026.


Tier 1: The kernel + attention crown (5)

Handle

Person

Why

@tri_dao

Tri Dao

FlashAttention 1/2/3 author, Mamba, Together AI. Frequent technical threads.

@cHHillee

Horace He

PyTorch core, torch.compile, FlexAttention. The “brrr” post author.

@jeremyphoward

Jeremy Howard

fast.ai, GPU MODE lectures, systems-first pedagogy.

@marksaroufim

Mark Saroufim

GPU MODE cofounder, PyTorch/Meta, TorchAO, TorchTitan.

@thom_wolf

Thomas Wolf

HuggingFace cofounder; deep systems posts, Ultra-Scale Playbook lineage.

Tier 2: Inference engines / research (7)

Handle

Person

Why

@woosuk_k

Woosuk Kwon

vLLM cofounder, PagedAttention lead author.

@zhuohan123

Zhuohan Li

vLLM cofounder, LMSYS.

@simon_mo_

Simon Mo

vLLM project lead.

@lm_zheng

Lianmin Zheng

SGLang cofounder, Chatbot Arena.

@ying11231

Ying Sheng

SGLang cofounder, RadixAttention.

@yzh119

Zihao Ye

FlashInfer author (CMU).

@HaoAILab

Hao AI Lab (Hao Zhang UCSD)

Speculative decoding, engine research.

Tier 3: Quantization + compression (4)

Handle

Person

Why

@Tim_Dettmers

Tim Dettmers

LLM.int8(), QLoRA, bitsandbytes, GPU-buying-guide author.

@younesbelkada

Younes Belkada

HF integrations for quant methods; bitsandbytes co-maintainer.

@gaunernst

Thien Tran

Quantized training, GPU MODE Lec 30.

@danielhanchen

Daniel Han

Unsloth author; low-level optimization posts.

Tier 4: Distributed / training systems (4)

Handle

Person

Why

@StasBekman

Stas Bekman

HF/Snowflake, distributed training practitioner writings.

@sea_snell

Charlie Snell

Test-time compute research; UCB.

@kellerjordan0

Keller Jordan

Modded-nanogpt speedrun, Muon optimizer, training microtricks.

@lambdaviking

William Merrill

Formal capabilities + parallelism research.

Tier 5: The pedagogues + synthesizers (5)

Handle

Person

Why

@karpathy

Andrej Karpathy

nanoGPT, education, occasional infra posts. Rate-limits himself, high signal.

@chipro

Chip Huyen

AI Engineering author; product/ML platform.

@rasbt

Sebastian Raschka

LLM From Scratch, monthly synthesis essays.

@simonw

Simon Willison

Daily curation + annotation of the entire LLM release firehose.

@natolambert

Nathan Lambert

RLHF/post-training + industry commentary; Interconnects.

Tier 6: Hardware + industry (3)

Handle

Person

Why

@dylan522p

Dylan Patel

SemiAnalysis; data center + GPU supply-chain intel.

@dzhulgakov

Dmytro Dzhulgakov

Fireworks AI CTO; production serving posts.

@soumithchintala

Soumith Chintala

PyTorch cofounder; industry perspective.

Tier 7: Guest passes (rotate in/out; 2 at a time) (2)

Use these slots for people relevant to what you’re currently building. Suggestions:

  • Building attention kernels? @tri_dao is already up; add @realDanFu (Dan Fu, Together / co-Mamba) and @simran_s_arora (ThunderKittens/Based).

  • DeepSeek/MoE deep-dive? DeepSeek doesn’t have canonical individual handles; follow the @deepseek_ai org account.

  • Reading vLLM source? Add @robertshaw21 (Robert Shaw, Neural Magic/RH), @mgoin_ (Michael Goin).

  • Local/GGUF? Add @ggerganov (Georgi Gerganov, llama.cpp).

  • Long-context/scaling? Add @_akhaliq (Ahmed Awadallah, HF paper firehose).

  • RL infra? Add @nrehiew_ / @vwxyzjn (Costa Huang, TRL).


Handles I intentionally did NOT include

  • Big-name AI CEOs and VCs. Signal-to-noise ratio catastrophic for this roadmap.

  • Random “AI influencer” accounts. They regurgitate primary sources with delay. Read the primary sources.

  • Anon accounts with no verifiable technical output. High risk, low reward.

  • Corporate accounts except @deepseek_ai and @vllm_project (which post announcements). Follow the people.


Weekly X protocol

  1. Open your List URL on Mon/Wed/Fri morning (15 min max).

  2. Save (bookmark) 3–5 threads/links per session.

  3. Batch-read the saved bookmarks Sunday morning during your r/LocalLLaMA half-hour.

  4. Reply only when you have a real technical addition. Never for engagement bait. Never “great post!” replies.

  5. Post-original-content cadence: maximum once every 2 weeks; content = your blog post / repo / benchmark. Not commentary.


When to post from your own account

Do post:

  • A concrete new artifact (blog, repo, benchmark) with a screenshot and a link. Once.

  • A well-scoped technical question a specific expert can answer, tagged politely. Once, and if no answer, drop it.

  • A correction to a wrong number in someone else’s post, backed by your reproduction. Rare, valuable, do it well.

Do NOT post:

  • Vague hot takes about the industry.

  • Reposts of press releases.

  • “Excited to share” without a link to the actual thing.

  • Anything late at night, angry, or under 200 characters of substance.


Handle-verification workflow

Before adding: check pinned tweet + bio + “About” panel for known-org affiliation. If you can’t verify identity in 30 seconds, don’t add. Impostor accounts of famous ML people exist.

If a handle above turns out inactive or renamed by the time you read this, search: [real name] site:x.com or [real name] twitter — pinned tweets and org affiliations are the disambiguator.