06 — Hardware Access Chokes At the Wrong Time

Failure class: Infrastructure risk. Modal timing: Months 3–7 (Phase 2 kernel work needs constant GPU access); month 8–12 for the multi-GPU experiments; scattered spikes for H100/H200 rentals. Aggregate probability of at least one significant hardware-access disruption: ~65%. Probability this alone kills the plan: ~15% — usually recoverable, but very expensive in time when it hits.


The Scenario

It is Aug 5, 2027. Your commit log has three specific gaps you remember viscerally: (a) three weeks in month 4 where the 3090 you were planning to buy fell through on OLX and you were stuck renting spot GPUs at 3x the expected cost, half-finishing kernels on a Colab session that kept timing out; (b) a month-long slowdown in month 6 when a RunPod pricing change turned your $80/mo rental budget into a $180 crunch and you cut experiments; (c) the week in month 9 you needed an H100 for FP8 experiments and couldn’t get one at any reasonable price for four days because a new frontier model release had spiked demand across every cloud.

None of these individually killed the plan. Together they cost you six weeks of momentum and about $600 of unbudgeted spend. The real damage was psychological: hardware-friction sessions are demoralizing in a way that pure difficulty is not. When the code is hard, you learn. When the infra is broken, you rage.


The Specific Failure Scenarios

6.1 You can’t get an H100 when you need it (probability ~50%)

Phase 3 (FA3 reading, FP8 experiments) and Phase 6 (TP scaling, disaggregated serving experiments) need modern datacenter GPUs — H100/H200 or MI300X. Availability is spiky. A big model release, an earnings-driven cloud-buy pause, or just a Tuesday afternoon can make on-demand H100s unavailable at reasonable prices for days.

Parry:

  • Batch H100 work into concentrated weekends. Do not need H100 daily. Identify the specific experiments that require Hopper features (TMA, wgmma, FP8) and schedule them as 1–2 dedicated H100 weekends per phase. Everything else (SGEMM ladder, Triton FA2 forward, RMSNorm, dequant kernels) runs on your 3090 or on cheaper A100 rentals.

  • Multi-provider accounts pre-created and prepaid. Sign up for RunPod, Lambda, Vast, and Modal in month 1. Add $20 credits to each. When you need an H100 hour on a Saturday morning, you do not want to also be filling out signup forms and waiting for KYC approval — you want to spin up on whichever provider has capacity, right now.

  • Watch for spikes: the days after a big frontier model release (GPT/Claude/Gemini/Llama) are the worst for GPU availability. Plan H100 work not those weeks.

6.2 Your 3090 dies mid-plan (probability ~15%)

Used 3090s are used. Fan bearings, VRAM thermal issues, PSU flakiness, or just an OLX unit that arrived already tired. If it dies in month 5, you are on rentals-only for a month while you source a replacement.

Parry:

  • Buy the 3090 in month 1, not month 3. The seed doc’s hardware section says “buy a used 3090” — do this first, before Phase 2 begins. Test it hard in the first week with a burn-in (Furmark, then a sustained 6-hour tensor-core load). If it’s going to die, you want it to die under warranty / return window, not in month 5.

  • A cheap $30 UPS protects against power flakiness that kills consumer GPUs over months.

  • Cloud fallback ready. The multi-provider account setup above doubles as your “3090 died” contingency.

6.3 Rental pricing spikes (probability ~40%)

RunPod/Vast pricing has been slowly rising. A sustained 2x spike would break the $50–150/mo budget. Global GPU demand is not going down.

Parry:

  • Own the boring 90%. Almost all of Phase 2’s kernel work runs on your local 3090. Rentals are for specific experiments (multi-GPU, H100 features, large-model inference). The more of the plan runs locally, the less pricing volatility hurts you.

  • Have a $250/mo escape budget. If rental prices jump, don’t fight it in the moment — pay it and adjust. Cutting experiments to save $60 is often a false economy because the lost time is worth more.

  • Also: know the “cheap” providers. Vast.ai is chaos but frequently 30–50% cheaper than RunPod. Community clouds like TensorDock or Prime Intellect can be even cheaper. Learn which providers to check when prices spike.

6.4 Colab / free-tier crutch collapses (probability ~30%)

Some early kernel work can be done on Colab or Kaggle. Both have been progressively tightening free-tier access. If you were relying on that for Phase 0 or Phase 1, plan for it to get worse, not better.

Parry: Do not build the plan on any free tier. Use them for one-off experiments only. Real work runs on hardware you control (your 3090) or hardware you rent explicitly.

6.5 The Zoho-workstation temptation (probability ~20%)

You might be tempted to use a Zoho-provisioned GPU workstation for some experiments. Don’t. It creates IP-attribution ambiguity for anything you eventually want to open-source, and it ties your learning to a machine you don’t own. Keep the personal and work rigs separate. This is a legal/career hygiene point, not a technical one — but it is important.

6.6 Power / thermal / space (probability ~25% for at least one incident)

A 3090 pulls 350W. In a Tamil Nadu summer, in a home office without good ventilation, the machine will throttle or shut down under sustained load. Your kernel benchmarks become garbage because clocks fluctuate.

Parry:

  • Ventilate the room. A basic tower fan aimed at the case is often enough.

  • Undervolt the 3090 — a well-undervolted 3090 gives up ~5% perf for ~25% lower power and much cooler operation. There are excellent guides on r/nvidia and r/overclocking. Do this in the first week.

  • Lock the GPU clocks (nvidia-smi -lgc <clock>) for benchmark runs so numbers are reproducible regardless of ambient temp. This is standard GPU-benchmark hygiene anyway.

  • A 1200 VA UPS is ~$80 in India and protects against the brownouts / outages that will otherwise cost you experiments.


The India-Specific Hardware Plan

This is where the plan needs to be honest about your geography. GPUs in India are expensive-plus-frustrating in ways non-Indians don’t fully appreciate.

Do NOT do

  • Do not import a GPU. Customs duty (~28% + GST) makes personal import of new GPUs financially irrational. A $1000 4090 lands at ~₹1.6L+ delivered, more than domestic new pricing after markup.

  • Do not buy a new 4090 unless money is not a constraint. ₹2.4L+ in India for a card whose value in your learning plan is roughly the same as a ₹55–70K used 3090.

  • Do not rely on the “gaming rig at home” plan long-term if the home is not yours to modify — ventilation and power constraints will bite.

DO — the pragmatic India plan

Tier 1 (baseline, month 1): Buy a used RTX 3090 24GB from a domestic marketplace. Sources ranked by reliability:

  1. Domestic Reddit-India GPU/PC-parts communities (r/IndianGaming, r/pcmasterrace_India). Lower fraud rate than OLX, better technical vetting.

  2. Reputable PC-building shops in Chennai/Bangalore that sell used enterprise pulls or trade-ins.

  3. OLX / Facebook Marketplace — cheapest but highest fraud rate. Meet in person, test on the seller’s PC before paying, run FurMark for 15 minutes, check for repainted shrouds (a sign of resale flip).

Expected price: ₹50,000–₹75,000 for a used 3090 as of mid-2026. Budget ₹80K to give yourself room to walk away from bad units.

Tier 2 (rentals): Multi-provider strategy, all set up in month 1:

  • RunPod — best UX, main workhorse, ~$1.4–2.5/hr for H100.

  • Lambda — competitive H100 pricing, good reliability, good for longer jobs.

  • Vast.ai — cheap and chaotic, best for one-off experiments where you can tolerate a stopped pod.

  • Modal — serverless GPU, ideal for benchmarking scripts you run occasionally without wanting a persistent VM.

Prepay ~$25 on each. Total setup cost: ₹8K. Ongoing target: ₹4K–12K/month across all providers combined. Higher months when you rent H100 weekends.

Tier 3 (aspiration, don’t rush): In Phase 6 / 7, if the plan is going well, consider:

  • Dual 3090 with NVLink — ~₹1.2L total, the classic r/LocalLLaMA 70B-4bit rig, gets you a real tensor-parallel testbed on your desk. This is a mid-plan upgrade (month 8+), not a start-of-plan purchase.

  • Apple silicon Mac with ≥32GB unified memory (if you have one already or can justify one) as an MLX / unified-memory lab. Do not buy one for this plan; use one if you have one.

Budget summary

Item

Cost (INR)

When

Used 3090

55,000–75,000

Month 1

UPS (1200 VA)

6,000

Month 1

Room ventilation (fan)

2,000

Month 1

Multi-provider prepay

8,000

Month 1

Ongoing rentals

4,000–12,000/mo

Months 3–13

Total year-1

~1.4–2.5 lakh

13 months

That is realistic. That is also achievable on a Zoho ML-engineer salary in Chennai. This is not a rich-hobbyist plan; it is a working-professional plan.


Early Warning Signals

  • Your 3090 acquisition drifts past month 2 for any reason (find a way to close it).

  • Any month where rental spend > ₹15K without a specific planned experiment justifying it (indicates drift, not use).

  • Two consecutive weekends where you wanted to run an experiment but couldn’t get a GPU at your price point.

  • Your GPU is running hot enough that clocks are dropping mid-run (check with nvidia-smi dmon).

  • A benchmark you ran last week does not reproduce this week (thermal or clock instability).

  • You’ve been “planning to try H100 experiments” for more than two phases without doing them.


Decision Tree


The Batching Discipline

The single largest cost optimization: plan your H100 work into 1–2 dedicated weekends per phase, rather than renting ad-hoc.

Concrete example. In Phase 3, you need Hopper access for FA3 study + one FP8 kernel experiment. Instead of five separate 2-hour rentals across the phase (10 hours × $2.50 = $25, plus setup overhead each time and reduced focus), block a single Saturday from 8am to 8pm as H100 Day:

  • Pre-plan the exact experiments to run (list of scripts, expected numbers).

  • Prep everything locally the day before — Docker image, dependencies, data.

  • Rent one H100 pod for 12 hours (~$25–35 at consumer rates), run the whole batch.

  • Write up the results the following Sunday.

You get more done, cheaper, with less context-switching overhead, and it fits cleanly into your calendar as a recurring “H100 Day per phase” ritual. This is the same principle as the “curiosity tax slot” in file 04 — bounded, scheduled, high-focus time boxes beat scattered dribbling.


Escalation Triggers

If you have not acquired the 3090 by end of month 2, this becomes a P0 issue. The Phase 2 kernel work is nearly impossible on rentals alone (constant iteration, latency-sensitive) and you cannot proceed effectively without local hardware. Actions:

  1. Escalate the search: post in r/IndianGaming, ask contacts at Zoho, check with local PC shops directly.

  2. Raise the budget ceiling to ₹90K if needed — the ₹15K delta is trivial next to a month of Phase 2 delay.

  3. If truly stuck (rare, but possible), rent a dedicated pod on RunPod for a month (~$150–200) to preserve tempo while continuing to hunt.

If sustained rental spend exceeds ₹18K/month for two consecutive months without exceptional experiments justifying it, do a hardware-cost review:

  • Are you renting GPUs to run experiments that would work fine on the 3090?

  • Are you leaving pods running (a classic cost bleed — check every provider’s dashboard)?

  • Are you renting bigger GPUs than needed for the experiment (A100 40GB when a 3090 24GB would suffice)?


The Bottom Line

Hardware access is a solved problem for a disciplined engineer with a modest budget and a plan. The failure mode is not “I couldn’t afford a GPU”; it is “I didn’t set up my hardware substrate early, and every time an experiment came up I lost 4 hours to friction.” Solve this in month 1. Buy the 3090. Set up the four rental accounts. Undervolt. Get a UPS. Lock clocks. Do H100 Day batching. Then never think about hardware again for the rest of the plan.

Every hour spent on hardware setup in month 1 saves ten hours of friction across months 2–13. It is one of the highest-ROI moves in the entire plan.