10 — GPU Procurement Literacy (2026)

Procurement is a technical skill. The person in the room who can say “H200 is the right buy for that workload, not B200, because your context length caps at 32k and you don’t need FP4” is the person the CFO listens to. Zoho’s on-prem play means you will be that person. This document is the cheat sheet.

Prices below are Q2–Q3 2026 snapshots (see Verification notes). GPU pricing is volatile: rental rates on Blackwell moved 3x in six months during supply constraints. Always re-check before quoting a customer.


1. The hierarchy in one table

GPU

Memory

Bandwidth

Peak dense

Rent low/high $/hr

Buy street $

Role

H100 SXM 80GB

80 GB HBM3

3.35 TB/s

989 BF16 TF

$1.49 – $6.98

$25–30K

Workhorse of production LLM inference. Default recommendation for 70B–120B.

H100 PCIe 80GB

80 GB HBM3

2.0 TB/s

756 BF16 TF

$1.99 – $4.50

$22–27K

Cheaper form factor, no NVLink between pairs of cards on many boards. Avoid for TP.

H200 SXM 141GB

141 GB HBM3e

4.8 TB/s

989 BF16 TF

$2.37 – $10.60

$28–35K

H100’s memory-fatter twin. Same compute, 43% more mem capacity + 43% more bandwidth. Best 70B single-GPU option; single-card 70B fp8 fits + real KV budget.

B200 SXM 192GB

192 GB HBM3e

8.0 TB/s

2.25 PF FP8

$2.25 – $16.11

$30–40K

Blackwell. Dense compute jump. Real workhorse when supply catches up in H2 2026. FP4 tensor cores.

B300 (Blackwell Ultra) 288GB

288 GB HBM3e

8.0 TB/s

~2.7 PF FP8

$2.45 (spot) – $6.80

~$45K est

Liquid cooling mandatory — most datacenters can’t take it. Highest capacity-per-GPU on market. On Llama-70B: ~100k tok/s FP8, ~150k tok/s FP4.

GB200 NVL72

72×192GB in one rack

130 TB/s aggregate

720 PF FP8 rack

$10.50 – $27.04 per GPU-hr

Rack-only, ~$3M

The trillion-token-model rack. Buy the frontier-scale inference is here. Serves DeepSeek-R1-class MoEs at wide-EP with room to spare.

L40S

48 GB GDDR6

864 GB/s

362 BF16 TF

$0.48 – $7.58 (typ ~$1)

$8–11K

Sweet-spot inference GPU for 7B–14B fp8 or 30B-4bit. No NVLink — bad for TP >2. Best per-dollar for CRM-scale agent workloads.

RTX Pro 6000 Blackwell 96GB

96 GB GDDR7

1.79 TB/s

~700 TF FP8

$0.23 – $17.27

$8–10K

Blackwell workstation card. 96GB GDDR7 changes the math — single card serves 70B fp8. Not datacenter-warranted; some clouds ban it.

RTX 6000 Ada 48GB

48 GB GDDR6

960 GB/s

362 BF16 TF

$0.14 – $1.57

$6–7K

Workstation-tier Ada. Similar to L40S, no NVLink.

RTX 5090 32GB

32 GB GDDR7

1.79 TB/s

~450 TF FP8

$0.18 – $0.99

$2.5–3.5K

Consumer. Datacenter TOS forbid in most clouds. Personal lab + edge only.

RTX 4090 24GB

24 GB GDDR6X

1.01 TB/s

330 BF16 TF

$0.08 – $1.61

$1.6–2.5K

Consumer classic. 24GB caps you at 8B-fp16 or 34B-4bit. r/LocalLLaMA workhorse.

RTX 3090 24GB

24 GB GDDR6X

936 GB/s

142 BF16 TF

$0.10 – $0.60

$600–800 used

Best kernel-learning GPU dollar for dollar. Buy one used.

AMD MI300X

192 GB HBM3

5.3 TB/s

1.3 PF FP8

$0.95 – $7.86 (Vultr $1.85)

$10–15K

The memory monster. 192GB single-GPU means 70B fp8 with 120GB KV budget or 235B MoE fp8 without TP. ROCm/vLLM 2026 support is real. Kernel maturity 10–25% behind CUDA on some ops.

AMD MI325X

256 GB HBM3e

6.0 TB/s

1.3 PF FP8

Estimate $1.50 – $6.00

~$12K–15K

HBM3e refresh. Character.AI moved Qwen3-235B FP8 here in 2026 and got 2x QPS on DigitalOcean. Multi-year 8-figure deal. Real for on-prem.

AMD MI355X

288 GB HBM3e

8.0 TB/s

2.5 PF FP8

Announced 2026

TBD

AMD’s Blackwell-class answer. Watch for enterprise deals; procurement teams should include AMD in RFPs.

Verification: rental price ranges are the min/max seen across 20+ providers on GetDeploying, Spheron, Vast, Lambda, RunPod, CoreWeave, Vultr, DigitalOcean, TensorWave, Crusoe (Q2 2026 snapshots). Purchase prices are US street prices from 2026 secondary-market and OEM channels.


2. Derived comparison: $/GB HBM and $/TFLOP FP8

The two ratios that actually determine fit.

GPU

$/GB HBM (buy)

$/GB HBM (rent-hr)

$/TFLOP FP8 (rent-hr)

Bandwidth/$-hr

H100 SXM 80GB

$325

$0.031

$0.0025

1340 GB/s per $

H200 SXM 141GB

$220

$0.020

$0.0025

2020 GB/s per $

B200 192GB

$180

$0.026

$0.0022

3160 GB/s per $

B300 288GB

$155 est

$0.017

Best-in-class

~2600 GB/s per $ (dedicated)

L40S 48GB

$200

$0.021

$0.0028

860 GB/s per $

RTX Pro 6000 96GB

$95

$0.015 (mid)

$0.0025

3600 GB/s per $

MI300X 192GB

$65

$0.010

$0.0014

2860 GB/s per $

MI325X 256GB

$50

$0.010 (est)

$0.0014

3000 GB/s per $

Read the table: AMD wins on $/GB HBM and $/TFLOP by a wide margin in 2026. Nvidia wins on kernel maturity, ecosystem, and — critically — resale/availability. The gap is closing every quarter; if your workload is bandwidth-limited (decode-heavy interactive chat), MI300X/MI325X should be in every RFP. If your workload is compute-limited (batch prefill, training), the ecosystem premium of Nvidia often still pays.

RTX Pro 6000 is the sleeper. 96GB GDDR7 at $8–10K buys you what H200 at $30K buys you in raw fit for 70B fp8. Two catches: not datacenter-warranted (Nvidia’s Enterprise SLA does not cover it), and no NVLink means no clean TP. For single-replica per card deployments in a lab or a Zoho on-prem edge box, this is dramatically undervalued.


3. Positioning by workload

The RFP-answer table. Memorize.

Workload

Right buy

Wrong buy

Why

7B–14B interactive chat, dense

L40S (48GB) or RTX Pro 6000 (96GB)

H100 (80GB)

You are paying 4x per GPU-hour for capacity you can’t use. L40S handles 60+ concurrent Qwen 14B fp8 easily.

30B–34B dense chat

H100 80GB single-card (fp8/AWQ) or RTX Pro 6000

2×L40S with TP

TP-2 over PCIe (no NVLink) kills throughput 30%+. Buy one bigger card.

70B dense chat, latency-sensitive

H200 141GB single-card fp8

H100 TP=4 for TTFT-hostile prefill

H200 removes the TP tax; single-card 70B fp8 fits with room for KV. If you already own H100s, TP=4 works — just budget for lower goodput.

70B dense, capacity-max

MI300X 192GB or B200 192GB

H100 TP=8 or H200 TP=2

Single-GPU 70B fp8 fits with 120GB KV budget — you get 6-8x concurrency of a TP=4 H100 replica.

MoE 235B (Qwen3-235B, Mixtral 8x22B)

MI325X 256GB single-node (following Character.AI’s playbook) or 2×B200

8×H100 TP=8

MoE needs capacity, not bandwidth-per-active-param. MI325X won this bake-off publicly.

MoE 400B+ (DeepSeek-V3/R1)

GB200 NVL72 with expert-parallelism

Anything smaller

The trillion-parameter models need the rack-scale product. Serve or don’t.

Batch summarization, non-interactive

L40S fleet or MI300X

H100/H200

Bandwidth doesn’t matter — throughput per dollar does. L40S on spot is the play.

Full FT of 70B / 100B pretrain

H100/H200/B200 with NVLink + IB

Consumer or PCIe-only

Comms bandwidth dominates training, not just per-card capacity.

Kernel dev / phase-2 learning

Used RTX 3090 (24GB)

Renting H100 by the month

Same PMPP/Triton exercises, 20x cheaper. Rent H100 only when you need TMA/wgmma/fp8.

Local r/LocalLLaMA lab

4090 or 5090 or dual 3090 with NVLink

Datacenter cards at home

Power, cooling, resale — all wrong for consumer environments.

Zoho angle. Zoho customers span the whole spectrum: SMB CRM instances that want a small chatbot (Qwen 14B fp8 on one L40S) all the way up to regulated-industry enterprises that want a private frontier-class deployment. Own this table. When Zoho sales gets asked “what hardware do we need to run your agentic feature private” and you can answer in five minutes with three defensible options at three price points, you become the person on the sales call.


4. Purchase vs rent decision matrix

Situation

Choice

Reasoning

POC / exploration / <3 months

Rent, cheapest neocloud

Volatility is fine; upside of not being locked.

Sustained inference, >6 months, >40% duty cycle

Buy (or 3-year reserved from neocloud)

Amortization crosses ~14 months. See §5 below.

Customer-specific on-prem deployment

Customer buys, Zoho specifies + operates

Zoho does not own customer hardware; you provide the sizing memo.

Bursty batch workload

Spot instances on multiple providers

Spot is 3-5x cheaper; batch jobs are restart-tolerant.

Interactive with strict SLO

Reserved + on-demand mix

Reserved for baseline (60–70% of capacity), on-demand for peaks. Never all-on-demand at scale.

Kernel/research lab (personal)

Buy used 3090 or 4090

Every day the GPU is not producing your learning, it’s costing you time not money.

B300/GB200 needs

Rent

You are not building a liquid-cooled datacenter in 2026.

Purchase amortization napkin

H100 SXM at $28K, 3-year straight-line, adds ~$1.07/hr per GPU. Add power (~$0.30/hr at $0.10/kWh, 700W avg), cooling (~$0.15/hr), datacenter space + networking + ops overhead (~$0.30/hr). All-in owned cost ≈ $1.80/hr vs $2.50/hr blended rent. Break-even ≈ 14 months at 100% duty, or ~22 months at 60% duty. Below 40% duty cycle, rent wins by default.

For MI300X: purchase $12K, 3-year = ~$0.46/hr + $0.75 overhead = ~$1.20/hr owned vs $1.85/hr rent. Break-even ~11 months. AMD’s TCO story is genuinely better if the software stack works for you.


5. Availability, in 2026, honestly

  • H100 SXM: available everywhere. Supply crisis of 2023–2024 is over. Prices dropped 30–50% during 2025. This is the boring, safe choice.

  • H200 SXM: widely available on specialist clouds. Hyperscalers charge 2-3x specialist pricing.

  • B200: available but volatile. March 2026 mean $5.09/hr, some providers spiked to $6+ during launch demand. Supply is stabilizing.

  • B300: limited. Liquid cooling requirement excludes ~80% of datacenters. Available on Spheron (spot $2.45), CoreWeave, and a few others.

  • GB200 NVL72: reserved capacity, waitlist through 2026. Frontier labs and hyperscalers get first allocation.

  • MI300X: wide availability, Vultr/TensorWave/Crusoe/Oracle/DigitalOcean. AMD is aggressive on OEM discounts for enterprise on-prem — get quotes.

  • MI325X: in production at DigitalOcean (Character.AI deal is the flagship customer), spreading through 2026.

  • L40S: ubiquitous. Every cloud has instant capacity.

  • RTX Pro 6000: limited on datacenter clouds due to TOS restrictions; Vast.ai and RunPod have it. For on-prem, order direct from Nvidia partners.

  • Consumer (5090/4090/3090): street availability normal. 5090 launch premium mostly gone by mid-2026.

Availability gotcha for on-prem. Lead times for H200/B200 through OEM channels (Dell, Supermicro, HPE) are 12–20 weeks in 2026 for standard SKUs, longer for custom configs. If a Zoho customer signs a contract in January requiring on-prem deployment, tell sales the delivery target is Q3, not next month. Building this expectation upfront is your job.


6. Consumer-tier reality (r/LocalLLaMA context)

Given your community credibility ambitions, be fluent in the enthusiast tier:

  • RTX 3090 (24GB, 936 GB/s): the enthusiast-lab benchmark. Two of them NVLinked run Llama-3-70B AWQ at ~15-25 tok/s decode. Total ~$1200 used.

  • RTX 4090 (24GB, 1008 GB/s): ~30% faster than 3090 on decode, no NVLink support on the card (Nvidia removed it), and a nightmare for pair setups.

  • RTX 5090 (32GB, 1.79 TB/s): Blackwell consumer. The 32GB opens up 20B-class fp16 comfortably and 70B-4bit tightly. FP8 tensor cores are the interesting bit.

  • Apple Mx Ultra 128GB/192GB unified memory: the other half of r/LocalLLaMA. Unified memory means CPU/GPU share; bandwidth is ~800 GB/s on M2 Ultra, ~1.1 TB/s on M4 Ultra. Runs 70B-4bit at 8-12 tok/s. MLX makes it usable.

Consumer-tier is where your public credibility gets built (benchmark posts, quant bake-offs). Nothing about consumer hardware is directly relevant to a Zoho on-prem deployment, but the ability to say “I ran this in my basement and here’s the numbers” is what separates you from the resume-only crowd.


7. The awkward geopolitics of GPU procurement

Not optional to think about. Named because ignoring it is unprofessional:

  • Export controls: H100/H200/B200 sales to some jurisdictions are restricted; on-prem deployments in those regions may require alternate SKUs (H20, etc.) or AMD/domestic alternatives. Zoho’s global footprint means this affects you.

  • Datacenter power availability: in some markets (US Northeast, parts of Europe), the constraint is not GPU supply, it is grid interconnection. B300/GB200 racks at 120–140kW each are a datacenter-planning conversation, not a purchase order.

  • Water and cooling: liquid-cooled Blackwell needs facility water. Air-cooled datacenters cannot host B300. If a Zoho customer’s DC is air-cooled and they want B300, the answer is “not without a retrofit.”

  • AMD vs Nvidia procurement politics: some enterprise customers have anti-Nvidia mandates (concentration risk, prior negotiations). Being able to spec an equivalent AMD stack is a genuine differentiator on those RFPs.


8. Reading list

  • Semi-Analysis GPU pricing reports (Q1, Q2, Q3 2026) — the reference for market data.

  • Character.AI × DigitalOcean AMD case study: https://www.digitalocean.com/blog/technical-deep-dive-character-ai-amd (the MI325X exemplar).

  • Nvidia H200 datasheet + Blackwell architecture whitepaper.

  • AMD MI300X and MI325X datasheets; ROCm 6.x/7.x release notes.

  • GetDeploying, Spheron, Vast, Lambda pricing pages — sample every quarter, keep a spreadsheet.

  • Chip Huyen’s AI Engineering (procurement chapter).


9. Exit test

You can produce, in one page:

  1. A defensible hardware recommendation (make, model, count, quantity, memory config, cooling constraint, purchase-vs-rent decision) for three customer profiles: (a) SMB CRM add-on chatbot serving 1000 employees, (b) mid-market regulated-industry agentic assistant serving 5000 employees, (c) enterprise multi-region CRM AI serving 100000+ employees. Include price per GPU and total capex or 12-month opex.

  2. The $/GB-HBM and $/TFLOP tables above from memory (approximately — you should know AMD wins these ratios in 2026, and by roughly how much).

  3. A clear articulation of when RTX Pro 6000 96GB is the right buy and when it is disqualified (datacenter warranty, cloud TOS).

  4. The names of at least three AMD-friendly clouds (Vultr, TensorWave, Crusoe, DigitalOcean, Oracle) and one credible enterprise reference customer (Character.AI on MI325X).

  5. The two-line answer to “should we buy or rent H200 for this project?” with the utilization threshold that flips the decision.