Issue #003 · NebulaPeer · GPU Infrastructure Intelligence for AI Startup Founders
Welcome back.
Over the past two issues we've established the problem — most AI startups default to the most expensive GPU option without ever questioning it — and the framework for thinking about the market in three tiers.
This week we go hands-on. The GPU cloud market in 2026 has expanded dramatically. There are now over 62 tracked providers offering nearly 3,000 distinct GPU configurations. The choice is no longer "which hyperscaler" — it's navigating a genuinely complex market with wildly different price points, reliability profiles, and use case fits.
This issue is your map.
The state of the market in 2026
The GPU cloud market is in the middle of a significant shift. A few things worth knowing before we get into the provider breakdown:
Supply has improved meaningfully. The GPU scarcity that dominated 2023 and 2024 has eased considerably. Supply chain improvements throughout 2025 increased availability across the board, and competition among cloud providers and specialised GPU services continues to intensify — providers now differentiate through more than price alone.
Prices have come down significantly. H100 rental prices have dropped from historical peaks of around $8/hour to a more reasonable range across most providers. New entrants and increased competition are driving continued compression, particularly at the Tier 2 and Tier 3 level.
New hardware is reshaping the hierarchy. Google Cloud made their latest B200 and GB200 instances generally available, competing directly with AWS, Azure, and Oracle Cloud offerings. The Blackwell generation is now live across multiple providers, and its arrival is pushing H100 and A100 pricing down further as demand shifts toward newer hardware.
The market is large and growing fast. The GPU rental market reached $7.38 billion in 2026 and is expected to grow another 28% in 2027. More money in the market means more providers, more competition, and more options for founders who know where to look.
The practical implication of all of this: the gap between Tier 1 and Tier 3 pricing is, if anything, wider than ever. More competition at the bottom of the market has driven Tier 3 prices down further while hyperscaler pricing has remained relatively sticky.
Tier 1 — Hyperscalers
AWS · Google Cloud · Microsoft Azure · Oracle Cloud
The hyperscalers are the market's starting point for most AI startups — and for a small subset of them, the right answer. Hyperscalers offer enterprise-grade security, compliance certifications, global region coverage spanning 60+ data centre locations, and deep integration with storage, networking, and managed ML services.
What they are not is cheap. Hyperscalers typically price 3–6x above specialist neoclouds for the same GPU because rented capacity comes bundled with enterprise SLA, compliance certifications, and cross-service integration. Specialised GPU providers are 60–85% cheaper than AWS, GCP, or Azure.
When Tier 1 is the right answer:
You have genuine enterprise compliance requirements — SOC 2, HIPAA, FedRAMP, ISO 27001
Your workload is deeply integrated with other cloud services that would be expensive to migrate
You have an existing committed spend agreement that makes GPU compute effectively discounted
Your customers contractually require hyperscaler infrastructure
When Tier 1 is the wrong answer:
You're early-stage with no compliance requirements
Your GPU workload is primarily training or batch processing
You've never actually compared the price of your workload elsewhere
Tier 2 — Specialist AI Clouds
CoreWeave · Lambda Labs · Crusoe Energy · Voltage Park · Northflank
GPU-first providers are purpose-built for AI and compute workloads, offering lower prices, faster GPU availability, and ML-specific features like pre-installed frameworks and Jupyter notebook access.
This tier has matured significantly in 2026. CoreWeave in particular has become a genuine enterprise alternative to AWS for AI workloads — with high-bandwidth networking for multi-node training, strong availability, and pricing meaningfully below hyperscaler rates. Lambda Labs continues to be a strong option for on-demand access with a clean onboarding experience. The sweet spot for most AI teams is an A100 80GB on a managed provider at around $1.50–$2/hour.
When Tier 2 is the right answer:
You need reliability and uptime guarantees but can't justify hyperscaler pricing
You're running large-scale multi-GPU training where networking quality matters
You need ML-native tooling and don't want to configure environments from scratch
Your monthly spend is $2,000–$10,000 and you want a single accountable provider
What to watch in Tier 2: Compliance certifications are improving but still catching up to hyperscalers. If compliance is a near-term requirement, check specific certifications before committing. Most Tier 2 providers now offer SOC 2 at minimum.
Tier 3 — GPU Marketplaces
Vast.ai · RunPod · TensorDock · Salad · Spheron
Peer-to-peer marketplaces aggregate GPU capacity from distributed sources, offering the lowest prices but with variable hardware quality and fewer enterprise guarantees.
The pricing in this tier is striking. Vast.ai marketplace hosts go as low as $0.31/hour for an RTX 4090. RunPod Community pricing for the same GPU sits around $0.69/hour. For H100 access, specialised providers are offering H100 access at as low as $1.38/hour — a fraction of what the same hardware costs on AWS.
For non-production workloads that tolerate marketplace host variability, Vast.ai wins on hourly cost. The key phrase is "tolerate variability" — this tier is not for production inference where your users are waiting. It is the right home for training runs, fine-tuning jobs, experiments, and any batch workload that has checkpointing set up.
When Tier 3 is the right answer:
Training, fine-tuning, and experimentation
Batch processing and data pipelines
Any workload that can checkpoint and resume
Cost-sensitive early-stage teams who want maximum compute per dollar
Testing and benchmarking before committing to a provider
When Tier 3 is the wrong answer:
Production inference with uptime SLAs
Compliance-sensitive workloads
Always-on services where interruption means user-facing downtime
The new entrants worth knowing about
Serverless GPU platforms — Modal, RunPod Serverless, Banana
A category that has grown significantly in 2026. Modal and RunPod Serverless both bill per-second of active compute, which can undercut hourly billing for sub-minute request workloads. The break-even is roughly 30% GPU utilisation: below that, serverless wins on total cost. Above that, dedicated hourly billing is cheaper because there is no cold-start overhead.
If your AI workload has variable or unpredictable traffic — common for early-stage products where you don't yet know your usage patterns — serverless GPU is worth evaluating seriously. You only pay when you're actually using compute.
Oracle Cloud
Increasingly competitive on H100 pricing and worth including in any serious provider comparison. Less well-known in the startup community but has been aggressive on GPU availability and pricing in 2026.
Voltage Park
A newer entrant that has made a significant amount of H100 capacity available at competitive rates. Worth a look for teams that have struggled with availability on other providers.
The hardware hierarchy in 2026
The GPU landscape isn't just about providers — it's about matching hardware to workload. Here's where things stand:
Blackwell B200 / GB200 — the newest generation. NVIDIA's Blackwell B200 series has emerged as a game-changer, delivering 10x throughput per megawatt compared to the previous generation. Available on major hyperscalers and some Tier 2 providers. Prices are still settling but trending down as supply increases. Worth evaluating for large-scale training where throughput per dollar matters.
H100 — still the workhorse for serious training. H100 pricing currently ranges from around $1.03/hour at the low end to significantly higher on hyperscalers. Widely available across all tiers. The right choice for most production training workloads.
A100 — increasingly affordable as Blackwell demand pulls attention upward. As new models mature, demand gradually shifts, lowering the cost of earlier generations such as the A100 and H100. This dynamic makes high-performance GPUs more accessible to startups. Excellent value for fine-tuning and training at scale.
RTX 4090 — the cost-conscious founder's best friend. Vast.ai marketplace hosts go as low as $0.31/hour for this hardware. For inference workloads and smaller training runs, the 4090 delivers strong performance at a fraction of data-centre GPU pricing. Start here unless you've genuinely outgrown it.
The decision framework for 2026
Use this to find where your workload belongs today:
Step 1 — Compliance check Do you have contractual or regulatory requirements that mandate a specific provider or certification? → Yes: Tier 1 only. → No: continue.
Step 2 — Workload type Is this training/fine-tuning/batch, or production inference? → Training/batch: Tier 3 marketplace, spot pricing. → Production inference: depends on your reliability and latency requirements.
Step 3 — For inference — can you tolerate any downtime? → No (production, SLA-bound): Tier 2 specialist cloud, on-demand or reserved. → Yes (internal, experimental): Tier 3 marketplace is fine.
Step 4 — What's your monthly GPU spend? → Under $2,000/month: Tier 3 marketplace. Don't overcomplicate it. → $2,000-10,000/month: Tier 2 specialist cloud worth evaluating. → Over $10,000/month: negotiate directly with Tier 2 providers. Volume matters.
What the smart teams are doing in 2026
The most cost-efficient AI startups are not picking a single provider and staying there. They're running a split stack:
Training and fine-tuning on Tier 3 marketplaces (Vast.ai, RunPod) at spot prices
Production inference on Tier 2 specialist clouds (CoreWeave, Lambda) for reliability
Variable inference on serverless platforms (Modal) to avoid paying for idle capacity
This approach captures the cost benefits of each tier for the workloads best suited to it. It's more setup than defaulting to a single provider — but the saving is significant enough that the setup cost pays back quickly.
The one action for this week
Pick one training workload your team is currently running on a Tier 1 provider.
Price it on Vast.ai or RunPod for the same GPU spec. Take into account egress and storage costs — not just the hourly rate. Calculate the monthly saving at your current usage level.
That number is the business case for a migration conversation with your team.
For most AI startups doing this exercise for the first time, the number is significant enough to act on.
What's coming in Issue #004
Next week: founder stories. Real AI startup teams sharing the infrastructure decisions they made, the mistakes that cost them, and the exact steps they took to cut their bills. If you have a GPU story worth sharing — a surprise bill, a painful migration, a provider that let you down — reply to this email. The best ones will feature anonymously in next week's issue.
If this issue was useful, forward it to a founder who's building with AI. That's the most helpful thing you can do for NebulaPeer right now.
NebulaPeer is independent. No GPU provider funds our editorial. No affiliate relationship shapes our recommendations. All pricing referenced in this issue is sourced from publicly available provider data as of May 2026 and subject to change — always verify current rates directly with providers before making decisions.
Questions, pushback, or a GPU war story? Reply to this email.
