Twitter/X Follow List¶
30 handles. Curated. No filler.
The philosophy¶
Do NOT scroll Twitter’s main timeline. It will destroy your attention span and radicalize you politically. Instead:
Create a private List called “Inference Eng” and add the 30 handles below.
Access ONLY via that list URL. Bookmark it. Delete the main app.
15 min × 5 weekdays = 75 min/week. Scan for: paper releases, kernel tricks, hiring signals, engine features shipping.
Never post political takes. Never quote-tweet drama. Never subtweet. This is a professional channel.
X handles change; verify by the person’s known org / pinned tweet before following. Handles below are known-good as of 2025–2026.
Tier 1: The kernel + attention crown (5)¶
Handle |
Person |
Why |
|---|---|---|
@tri_dao |
Tri Dao |
FlashAttention 1/2/3 author, Mamba, Together AI. Frequent technical threads. |
@cHHillee |
Horace He |
PyTorch core, torch.compile, FlexAttention. The “brrr” post author. |
@jeremyphoward |
Jeremy Howard |
fast.ai, GPU MODE lectures, systems-first pedagogy. |
@marksaroufim |
Mark Saroufim |
GPU MODE cofounder, PyTorch/Meta, TorchAO, TorchTitan. |
@thom_wolf |
Thomas Wolf |
HuggingFace cofounder; deep systems posts, Ultra-Scale Playbook lineage. |
Tier 2: Inference engines / research (7)¶
Handle |
Person |
Why |
|---|---|---|
@woosuk_k |
Woosuk Kwon |
vLLM cofounder, PagedAttention lead author. |
@zhuohan123 |
Zhuohan Li |
vLLM cofounder, LMSYS. |
@simon_mo_ |
Simon Mo |
vLLM project lead. |
@lm_zheng |
Lianmin Zheng |
SGLang cofounder, Chatbot Arena. |
@ying11231 |
Ying Sheng |
SGLang cofounder, RadixAttention. |
@yzh119 |
Zihao Ye |
FlashInfer author (CMU). |
@HaoAILab |
Hao AI Lab (Hao Zhang UCSD) |
Speculative decoding, engine research. |
Tier 3: Quantization + compression (4)¶
Handle |
Person |
Why |
|---|---|---|
@Tim_Dettmers |
Tim Dettmers |
LLM.int8(), QLoRA, bitsandbytes, GPU-buying-guide author. |
@younesbelkada |
Younes Belkada |
HF integrations for quant methods; bitsandbytes co-maintainer. |
@gaunernst |
Thien Tran |
Quantized training, GPU MODE Lec 30. |
@danielhanchen |
Daniel Han |
Unsloth author; low-level optimization posts. |
Tier 4: Distributed / training systems (4)¶
Handle |
Person |
Why |
|---|---|---|
@StasBekman |
Stas Bekman |
HF/Snowflake, distributed training practitioner writings. |
@sea_snell |
Charlie Snell |
Test-time compute research; UCB. |
@kellerjordan0 |
Keller Jordan |
Modded-nanogpt speedrun, Muon optimizer, training microtricks. |
@lambdaviking |
William Merrill |
Formal capabilities + parallelism research. |
Tier 5: The pedagogues + synthesizers (5)¶
Handle |
Person |
Why |
|---|---|---|
@karpathy |
Andrej Karpathy |
nanoGPT, education, occasional infra posts. Rate-limits himself, high signal. |
@chipro |
Chip Huyen |
AI Engineering author; product/ML platform. |
@rasbt |
Sebastian Raschka |
LLM From Scratch, monthly synthesis essays. |
@simonw |
Simon Willison |
Daily curation + annotation of the entire LLM release firehose. |
@natolambert |
Nathan Lambert |
RLHF/post-training + industry commentary; Interconnects. |
Tier 6: Hardware + industry (3)¶
Handle |
Person |
Why |
|---|---|---|
@dylan522p |
Dylan Patel |
SemiAnalysis; data center + GPU supply-chain intel. |
@dzhulgakov |
Dmytro Dzhulgakov |
Fireworks AI CTO; production serving posts. |
@soumithchintala |
Soumith Chintala |
PyTorch cofounder; industry perspective. |
Tier 7: Guest passes (rotate in/out; 2 at a time) (2)¶
Use these slots for people relevant to what you’re currently building. Suggestions:
Building attention kernels? @tri_dao is already up; add @realDanFu (Dan Fu, Together / co-Mamba) and @simran_s_arora (ThunderKittens/Based).
DeepSeek/MoE deep-dive? DeepSeek doesn’t have canonical individual handles; follow the @deepseek_ai org account.
Reading vLLM source? Add @robertshaw21 (Robert Shaw, Neural Magic/RH), @mgoin_ (Michael Goin).
Local/GGUF? Add @ggerganov (Georgi Gerganov, llama.cpp).
Long-context/scaling? Add @_akhaliq (Ahmed Awadallah, HF paper firehose).
RL infra? Add @nrehiew_ / @vwxyzjn (Costa Huang, TRL).
Handles I intentionally did NOT include¶
Big-name AI CEOs and VCs. Signal-to-noise ratio catastrophic for this roadmap.
Random “AI influencer” accounts. They regurgitate primary sources with delay. Read the primary sources.
Anon accounts with no verifiable technical output. High risk, low reward.
Corporate accounts except @deepseek_ai and @vllm_project (which post announcements). Follow the people.
Weekly X protocol¶
Open your List URL on Mon/Wed/Fri morning (15 min max).
Save (bookmark) 3–5 threads/links per session.
Batch-read the saved bookmarks Sunday morning during your r/LocalLLaMA half-hour.
Reply only when you have a real technical addition. Never for engagement bait. Never “great post!” replies.
Post-original-content cadence: maximum once every 2 weeks; content = your blog post / repo / benchmark. Not commentary.
When to post from your own account¶
Do post:
A concrete new artifact (blog, repo, benchmark) with a screenshot and a link. Once.
A well-scoped technical question a specific expert can answer, tagged politely. Once, and if no answer, drop it.
A correction to a wrong number in someone else’s post, backed by your reproduction. Rare, valuable, do it well.
Do NOT post:
Vague hot takes about the industry.
Reposts of press releases.
“Excited to share” without a link to the actual thing.
Anything late at night, angry, or under 200 characters of substance.
Handle-verification workflow¶
Before adding: check pinned tweet + bio + “About” panel for known-org affiliation. If you can’t verify identity in 30 seconds, don’t add. Impostor accounts of famous ML people exist.
If a handle above turns out inactive or renamed by the time you read this, search: [real name] site:x.com or [real name] twitter — pinned tweets and org affiliations are the disambiguator.