Own GPU Choice: 3090 vs 4090 vs 5090 (2026 reality)¶
TL;DR decision matrix¶
Scenario |
Buy this |
|---|---|
First GPU, budget-limited ($700-900) |
Used RTX 3090 (24GB, ~$600-800 used) |
Want fastest single card, budget ~$1600 |
RTX 4090 (24GB, ~$1500-1800 used/refurb) |
Want to run 70B-4bit natively (48GB) |
Dual RTX 3090 with NVLink (~$1400-1600 total) |
Money is not the constraint |
RTX 5090 (32GB) or wait for used 4090 supply |
Already have a 3090, upgrading |
Add a second 3090 before switching to 4090 |
Strong recommendation for your situation: used 3090. Then in ~9 months, add a second 3090 for NVLink dual-24GB.
The three cards at a glance (2026 datapoints)¶
Spec |
RTX 3090 |
RTX 4090 |
RTX 5090 |
|---|---|---|---|
Arch |
Ampere (GA102) |
Ada Lovelace (AD102) |
Blackwell (GB202) |
CUDA compute cap |
sm_86 |
sm_89 |
sm_120 |
VRAM |
24 GB GDDR6X |
24 GB GDDR6X |
32 GB GDDR7 |
Mem BW |
936 GB/s |
1008 GB/s |
1792 GB/s |
FP16 tensor |
142 TFLOPs |
330 TFLOPs |
~450 TFLOPs |
BF16 tensor |
142 TFLOPs |
330 TFLOPs |
~450 TFLOPs |
FP8 tensor |
— (not supported) |
660 TFLOPs |
~900 TFLOPs |
FP4 tensor |
— |
— |
supported |
NVLink |
Yes (bridge) |
No (removed) |
No |
TDP |
350W |
450W |
575W |
PSU |
750W min |
1000W min |
1200W min |
Used 2026 street price (USD) |
$600-800 |
$1500-1800 |
$2000-2500 (new, scarce) |
All 3 = consumer GeForce cards. None can be legally used in commercial data centers per NVIDIA EULA. Nobody enforces this on personal machines. Don’t rack them in a corporate DC.
Why the used 3090 wins for you¶
1. 24GB is the magic threshold¶
At 24GB you can:
Run Llama-3.1-8B FP16 with generous KV cache.
Run Llama-3.3-70B-Q4_K_M (~40GB) with CPU offload — painful but possible.
Run Qwen3-32B-Q4 (~18GB) with 32k context KV cache.
Load full-precision 13B models.
Actually profile a real KV cache with meaningful memory pressure.
A 3080/4070 with 12GB cannot do any of this. You’ll spend more time swapping than learning.
2. NVLink is only on 3090 (and gone forever)¶
NVIDIA removed the NVLink bridge on the 4090 and 5090. This is a strategic decision to force enterprise buyers to A100/H100. If you ever want a real dual-GPU setup with fast peer-to-peer, the 3090 is the only consumer option. Two 3090s + NVLink bridge = 48GB effective VRAM with ~100GB/s inter-GPU bandwidth. Two 4090s = 48GB but forced through PCIe (~64GB/s, and much higher latency).
For tensor parallelism experiments, ring attention on 2 devices, or running 70B-4bit locally, dual 3090 is the hobbyist king.
3. Ampere is still modern enough¶
sm_86 supports:
cp.async(async global-to-shared copy)BF16 and FP16 tensor cores
WMMA (warp matrix multiply accumulate)
CUDA graphs
Modern nvcc / Triton / PyTorch
sm_86 lacks:
FP8 tensor cores (Hopper+ only)
TMA (Hopper+)
WGMMA (Hopper+)
FP4 (Blackwell only)
For Phase 2-3 (kernels + attention) work through 90% of Phase 6 (distributed), Ampere is fine. You’ll want to rent H100 time to touch WGMMA/TMA/FP8 — that’s what 02_rent_gpu_strategy.md is for.
4. Used market is real¶
3090s are ex-crypto or ex-gamer cards. In 2026 the market is mature. Buy from:
eBay — sort by “used, sold” to see comps. Look for 6-month seller history, US or EU shipping if you’re India.
/r/hardwareswap (US) or /r/IndianHardwareSwap.
Facebook Marketplace local pickup.
Micro Center / Craigslist (US only).
India: OLX, Nehru Place (Delhi), SP Road (Bangalore) shops — haggle. Expect ₹55,000-70,000 in Q4 2026.
Buying checklist:
Ask for the S/N and check on NVIDIA’s warranty portal.
Ask for a screenshot of GPU-Z showing hash rate history (crypto tell) — not a dealbreaker but negotiate down.
Confirm the seller will let you stress-test for 15 min before payment.
Run
nvidia-smi -q— check power limits, VRAM ECC errors, memory clock stability.Run a 15-min
gpu-burnor a Llama-3.1-8B benchmark. Watch temps: should stay under 80°C.
Red flags: melted 12VHPWR connector traces (rare on 3090; more common on 4090), missing shroud screws (opened for cleaning?), thermal-paste weep marks.
When to pick the 4090 instead¶
You already have a 3090 and want a second card of newer arch — unlikely a good move; better to get a matching 3090 for NVLink.
You want to touch FP8 tensor cores locally without renting. 4090 has FP8; 3090 doesn’t.
You value single-card throughput for training small models (e.g., you’ll do a 1B-param LM pretraining run at home). 4090 is ~2.3x faster than 3090 for tensor-core-heavy training.
You’re rich or your employer is paying.
4090 disadvantages: No NVLink. 450W TDP means a 1000W PSU and real thermals. 12VHPWR connector — seat it properly.
When to pick the 5090¶
You have >$2500 to spend.
You want a future-proof card that will last through the whole 13 months and beyond.
You want FP4 tensor cores — real 4-bit inference is where the field is going in 2026-27.
You have a 1200W PSU already, and case cooling to match.
Honestly, in 2026 the 5090 is still supply-constrained and priced at scalper levels. Wait 6 months, or just get the 3090.
Dual 3090 economics (the actual hobbyist play)¶
Component |
Cost |
|---|---|
2× used RTX 3090 |
~$1400 |
NVLink bridge (3-slot) |
~$100 (used) |
Motherboard with 2× PCIe 4.0 x16 (or x8/x8) |
~$200-300 (used AM4 X570 or LGA1200 Z590) |
CPU (Ryzen 5900X used or i5-13600K) |
~$200-300 |
RAM (64GB DDR4) |
~$120 |
PSU (1200W Platinum) |
~$180 |
Case (open air or Fractal Meshify) |
~$100 |
Cooling (2x AIO or beefy air) |
~$150 |
Storage (2TB NVMe) |
~$100 |
Total build |
~$2500-2800 |
Compare to renting: at $2.20/hr for a single H100, $2500 = ~1136 hours = ~47 days of 24×7 rental. You will get more done with 12 months of unlimited local time than with 47 days of stress-rental.
But: you cannot do WGMMA or FP8 on 3090s. You still need rented H100 hours for FA3/Hopper-specific work. Budget ~$300 total across the 13 months for rented burst work.
For your specific situation (Chennai/Bangalore, Zoho salary)¶
Right now: used 3090 from OLX / a trusted local shop. ₹55,000-70,000.
Month 6-8: if the roadmap is working and you’re on track, add a second 3090. Total investment ~₹1.4L.
Rent H100 time as needed (see
02_rent_gpu_strategy.md).Never buy a 4090 or 5090 with your own money for this roadmap. The delta over dual-3090 doesn’t justify the cost for 90% of what you’ll actually do. If you get a job that requires it, they’ll buy it.
Power and cooling reality¶
A single 3090 at full load draws 350W. Add 150W for CPU/board/fans = 500W system. A 750W PSU is fine.
Dual 3090 at full load = 700W GPU + 200W rest = 900W. 1200W PSU minimum, 1000W is too tight.
Ambient temperature in Indian summer: your GPU will run 5-8°C hotter. If your room hits 32°C, the GPU will hit 85°C easily. AC or open-air rig in a well-ventilated room. Not a laptop or a sealed cabinet.
Noise: Founders Edition 3090 is quieter. AIB (ASUS Strix, MSI Suprim) cards are louder but cooler. If your desk is your bedroom, prioritize acoustics.
Cross-references¶
Rental strategy for what a local 3090 can’t do:
02_rent_gpu_strategy.mdBudget planning:
04_budget_scenarios.mdH100/H200/B200 landscape for context:
05_2026_gpu_landscape.mdWhat kernel work needs H100 vs runs fine on 3090:
03_gpu_kernels/(owned by another swarm agent).
One last time: buy a used 3090, stop debating, start building.