2026 GPU Landscape (H100 / H200 / B200 / B300 / GB200 / MI300X / MI325X)¶
This is not about buying these. You will never own one. This is about understanding what your rented pod actually is, and speaking the language when an study partner says “what’s your take on GB200 NVL72 vs 8xH100 for our workload?”
NVIDIA Datacenter GPU family tree¶
GPU |
Arch |
Year |
HBM |
Mem BW |
FP16 TC |
FP8 TC |
FP4 TC |
TDP |
Rough MSRP |
|---|---|---|---|---|---|---|---|---|---|
A100 40GB |
Ampere |
2020 |
40 HBM2e |
1555 GB/s |
312 TF |
— |
— |
400W |
~$10K |
A100 80GB |
Ampere |
2021 |
80 HBM2e |
2039 GB/s |
312 TF |
— |
— |
400W |
~$15K |
H100 SXM 80GB |
Hopper |
2022 |
80 HBM3 |
3350 GB/s |
989 TF |
1979 TF |
— |
700W |
~$30K |
H100 PCIe 80GB |
Hopper |
2022 |
80 HBM3 |
2000 GB/s |
756 TF |
1513 TF |
— |
350W |
~$25K |
H200 SXM 141GB |
Hopper |
2024 |
141 HBM3e |
4800 GB/s |
989 TF |
1979 TF |
— |
700W |
~$35K |
B100 (SXM) 192GB |
Blackwell |
2024/25 |
192 HBM3e |
8000 GB/s |
~1800 TF |
~3500 TF |
~7000 TF |
700W |
~$35K |
B200 (SXM) 192GB |
Blackwell |
2025 |
192 HBM3e |
8000 GB/s |
2250 TF |
4500 TF |
9000 TF |
1000W |
~$40K |
B300 (Ultra) 288GB |
Blackwell Ultra |
2025-26 |
288 HBM3e |
~12 TB/s |
~2500 TF |
~5000 TF |
~10 PF |
1200W |
~$50K |
GB200 NVL72 |
Grace+Blackwell rack |
2025 |
72 GPUs × 192GB |
576 TB HBM total |
— |
— |
— |
120 kW/rack |
~$3M/rack |
Notes:
SXM vs PCIe: SXM is board-form-factor with NVLink between GPUs on the same board (fast). PCIe is card-form-factor, slower peer-to-peer. Most “H100” you rent from RunPod is SXM.
FP8 support debuted with Hopper. FP4 debuted with Blackwell (B200/B300).
Numbers above are peak dense tensor-core TFLOPs. Sparse (2:4) versions are 2x these.
“NVL72” = 72 Blackwell GPUs interconnected via 5th-gen NVLink switch, appears as one very-large-GPU to software. Aimed at trillion-param LLM serving.
AMD Instinct family (the real challenger in 2026)¶
GPU |
Year |
HBM |
Mem BW |
FP16 TC |
FP8 TC |
Rough MSRP |
|---|---|---|---|---|---|---|
MI250X |
2021 |
128 HBM2e |
3200 GB/s |
383 TF |
— |
~$15K |
MI300X |
2023 |
192 HBM3 |
5300 GB/s |
1300 TF |
2600 TF |
~$15K |
MI325X |
2024 |
256 HBM3e |
6000 GB/s |
1300 TF |
2600 TF |
~$18K |
MI355X |
2025 |
288 HBM3e |
~8 TB/s |
2500 TF |
5000 TF (+ FP4) |
~$25K |
MI400 |
2026 |
432 HBM4 |
~19 TB/s |
~5000 TF |
~10 PF (FP8/FP4) |
~$30K (announced) |
Why AMD matters in 2026:
Meta and Microsoft are running MI300X in real production for Llama-3-405B / GPT-scale inference. MI300X has 192GB HBM vs H100’s 80GB — fits a 405B-fp16 across 8 cards where H100 needs 16.
ROCm 7.x is finally usable for LLM inference (2025 release cycle). vLLM has first-class ROCm support. PyTorch upstream ROCm works.
Price/performance for memory-bound decode: MI300X is often 20-30% cheaper than H100 on inference-per-dollar for large models.
Rentability: RunPod and Prime Intellect list MI300X at ~$<phone_number_or_numberic_id_or_random_id_141>/hr. If you rent for one weekend you can genuinely write a “ROCm inference on MI300X” portfolio post.
MI400 in 2026: if the timeline holds, this is the H200-killer. Watch AMD’s Q3 2026 Financial Analyst Day for real numbers.
Positioning tables¶
$/GB HBM (2026 rental)¶
GPU |
HBM |
$/hr |
$/hr/GB |
|---|---|---|---|
H100 80GB |
80 GB |
$2.20 |
$0.028 |
H200 141GB |
141 GB |
$3.50 |
$0.025 |
B200 192GB |
192 GB |
$5.50 |
$0.029 |
MI300X 192GB |
192 GB |
$3.20 |
$0.017 |
MI300X is meaningfully cheaper per GB of HBM. For KV-cache-heavy workloads (long context, large batch), this is a real advantage.
$/TFLOP FP8 (2026 rental)¶
GPU |
FP8 TF |
$/hr |
$/TF |
|---|---|---|---|
H100 80GB |
1979 |
$2.20 |
$0.00111 |
H200 141GB |
1979 |
$3.50 |
$0.00177 |
B200 192GB |
4500 |
$5.50 |
$0.00122 |
MI300X 192GB |
2600 |
$3.20 |
$0.00123 |
Compute-per-dollar is roughly equivalent across the fleet. The real differentiator is memory bandwidth and capacity, not FLOPs.
Memory-bandwidth-per-dollar (the decode-throughput proxy)¶
GPU |
Mem BW GB/s |
$/hr |
GB/s/$ |
|---|---|---|---|
H100 80GB |
3350 |
$2.20 |
1523 |
H200 141GB |
4800 |
$3.50 |
1371 |
B200 192GB |
8000 |
$5.50 |
1455 |
MI300X 192GB |
5300 |
$3.20 |
1656 |
MI300X wins on memory bandwidth per dollar. For decode-bound inference (small batch, long generation), this is the dominant metric. This is why Meta uses MI300X.
Availability reality (2026)¶
H100: widely available now. Used H100 SXM boards trickling into secondary market but $10K+ still. No home use.
H200: available on all major clouds. Slightly harder to get 8x SXM.
B100/B200: hyperscaler-first. Available on AWS/GCP/Azure but with premium pricing. RunPod is starting to list them mid-2026.
B300 / GB200 NVL72: capacity-constrained. Only the largest labs and hyperscalers. Do not expect to rent this for at least another 12-18 months on retail clouds.
MI300X: widely rentable at RunPod, Vast, TensorWave (AMD-specialty). Cheaper than H100.
MI325X: limited but growing.
MI355X: just launched (2025), limited access via TensorWave, HotAisle, and some hyperscalers.
Interconnect landscape¶
Interconnect matters as much as raw GPU for large-model inference.
Fabric |
GPUs connected |
Bandwidth |
Where |
|---|---|---|---|
PCIe 5.0 x16 |
Any 2 GPUs |
63 GB/s bidir |
Consumer + workstation |
NVLink 4 |
H100 SXM (8-way) |
900 GB/s per GPU |
HGX H100 systems |
NVLink 5 |
Blackwell SXM (8-way) |
1800 GB/s per GPU |
HGX B200 |
NVSwitch (NVL72) |
72 Blackwell GPUs |
Rack-scale, ~130 TB/s aggregate |
GB200 NVL72 |
Infinity Fabric |
8x MI300X |
896 GB/s per GPU |
Instinct platforms |
InfiniBand NDR 400G |
Node-to-node |
400 Gb/s = 50 GB/s |
Real HPC clusters |
RoCE 400G |
Node-to-node |
400 Gb/s |
Cheaper alternative to IB |
Rule of thumb for tensor parallel: you need NVLink-class intra-node bandwidth or you’ll bottleneck. For pipeline parallel or ZeRO across nodes, InfiniBand 400G is table stakes.
Cloud vendor GPU inventory (approx, 2026)¶
Vendor |
H100 |
H200 |
B200 |
GB200 |
MI300X |
|---|---|---|---|---|---|
AWS |
Yes (p5.48xl) |
Limited |
Yes (p6) |
Preview |
No |
GCP |
Yes (a3-highgpu) |
Yes |
Yes (a4) |
Preview |
No |
Azure |
Yes (ND H100 v5) |
Yes |
Yes (ND B200 v6) |
No |
Preview |
Oracle Cloud |
Yes |
Yes |
Yes |
Yes |
Yes |
Coreweave |
Yes |
Yes |
Yes |
Yes |
No |
Lambda |
Yes |
Yes |
Yes |
No |
No |
RunPod |
Yes |
Yes |
Yes |
No |
Yes |
Vast |
Yes |
Some |
Emerging |
No |
Some |
TensorWave |
— |
— |
— |
— |
Yes (specialty) |
Prime Intellect |
Yes |
Yes |
Yes |
No |
Yes |
Oracle is the interesting one: aggressive on Blackwell/GB200 capacity contracts. Coreweave is second in NVIDIA-heavy inventory.
The next 18 months (2026-2027 forecast)¶
Stuff to watch, not to bet on:
NVIDIA Rubin (successor to Blackwell) — rumored 2026 announce, 2027 ship. Rubin Ultra with HBM4. This will make B200 the “H100 of the Blackwell era” — widely available, cheap on secondary market by 2028.
AMD MI400 — announced for 2026, HBM4, will likely be the first serious NVIDIA alternative in AI training.
Google TPU v5p / v6e — not directly rentable outside GCP but understand the TPU model for study conversations.
Groq, Cerebras, Etched, Tenstorrent — specialty inference silicon. Groq in particular is deployed at scale for fast inference. Don’t invest career time here yet, but be aware.
Intel Gaudi 3 / Gaudi 4 — exists, has some hyperscaler pickup, but a distant third to NVIDIA/AMD. Not worth learning specifically.
The through-line: in 2026, NVIDIA has ~85% market share for AI training + ~70% for inference. AMD is 10% and growing. Everyone else is <5%. Your career leverage is on NVIDIA CUDA + optional AMD ROCm. Ignore the rest.
The one study answer to have ready¶
Q: If you were sizing inference infrastructure for a 70B-scale model at 100k QPS, what GPU family would you pick and why?
A: 70B-fp16 is 140GB, so with KV cache and headroom you want ~192GB per model replica. Two options: 2x H100-80GB with tensor parallel and disaggregated prefill/decode, or 1x MI300X-192GB fitting the full model on one card and scaling out horizontally. TP-across-H100 gives lower TTFT but higher $/token; single-card MI300X gives simpler ops and better $/token but potentially worse tail latency. At 100k QPS I’d lean disaggregated H200 pods (141GB fits 70B fp16 + KV natively) with LMCache-style KV reuse across replicas. But I’d A/B against MI300X on real prod traffic before locking in, since decode is memory-bound and MI300X has 60% more bandwidth per dollar.
That is the answer. Memorize the shape, not the numbers.
Cross-references¶
Rental strategy per vendor:
02_rent_gpu_strategy.mdOwn GPU decision:
01_own_gpu_choice.mdDistributed papers where these platforms are discussed:
09_papers/05_distributed_papers.md(DeepSeek-V3 uses 2048 H800; Mooncake uses H100+H800).SemiAnalysis is the reference for real hardware economics:
10_communities/04_blog_canon.md.
One line: you rent NVIDIA H100/H200 for 95% of your work, rent MI300X once for a portfolio piece, and speak fluently about GB200 NVL72 for studies.