
Baseten
B2B usage-based cloud platform. Sells 'deployments' rather than pure tokens: dedicated per-minute autoscaling GPU deployments, plus token-priced Model APIs and training/embeddings. Revenue scales with customer inference volume; COGS is dominated by rented/committed GPU capacity.
Chronological priced rounds. Series C valuation (~$825M) is secondary-reported; the ~$5B Series E anchor is triangulated from the Series F re-rate language ($5B -> up to $13B in ~5 months). ~$2.085B raised across all rounds. Intermediate rounds may exist but are not separately price-confirmed here.
Earnings, margins, COGS & capex
Explosive top-line growth on inference demand: ARR from ~$200M (Dec 2025) to ~$600M annualized (Mar 2026), ~20x YoY. Platform reportedly serves >1 billion inference calls/day across 87 clusters on 18 clouds. Unit economics are GPU-cost-constrained and undisclosed; the company is optimizing model-serving throughput (speculative decoding, Blackwell-class hardware) to defend spread. No audited financials are public.
Income statement — where each revenue dollar goes
% of revenueOf every $1 of revenue, ~50¢ is cost of goods and ~50¢ operating expense, leaving ~0¢ of operating profit.
Revenue trend
Margins
structurally pressured by GPU COGS; company claims optimization (weight caching, container orchestration, workload distribution) aids it, but no figure released
assumed deeply negative (growth-stage)
COGS structure
Dominated by GPU compute rented/committed across hyperscalers and neoclouds (18 cloud environments). Token-priced APIs run a spread model (buy compute wholesale, sell inference retail); dedicated deployments pass GPU-hours through at per-minute pricing. Cost/token has been falling rapidly (widely-cited industry estimates of roughly an order of magnitude per year, with some benchmarks quoted at ~10-50x since 2023), which both compresses list prices and improves throughput economics.
Capex
Asset-light relative to a hyperscaler: Baseten orchestrates third-party GPU capacity rather than owning data centers, so 'capex' shows up as multi-cloud capacity commitments and GPU reservations, not owned fixed assets. Magnitude not disclosed.
Latest earnings
not applicable (private)
No formal guidance. Company narrative frames itself as foundational inference infrastructure for the 'next phase' of AI adoption.
- ARR (Mar 2026)
- ~$600M annualized
- YoY growth
- ~20x (~1,900%)
- Inference volume
- >1B calls/day
- Infrastructure footprint
- 87 clusters across 18 clouds
- Series F raise
- $1.5B (Jun 22 2026)
- Total raised
- ~$2.085B across all rounds
- Post-money valuation
- up to $13B (Jun 2026; $11B/$13B tranches)
Growth drivers
- Secular shift of AI spend from training to inference as models move into production
- Open-weight model wave (Llama, DeepSeek, Qwen, etc.) driving demand for third-party serving instead of closed APIs
- Marquee AI-native customers scaling inference (Cursor, Notion, Abridge, Clay)
- Multi-cloud sourcing (87 clusters / 18 clouds) enabling GPU supply arbitrage and capacity availability
- Performance edge (fast cold starts, autoscaling, speculative decoding, Blackwell-class hardware) lowering cost/token and latency
- Enterprise compliance (HIPAA, SOC 2) opening regulated verticals (healthcare via Abridge, etc.)
Bull & bear
Baseten is a picks-and-shovels leader in the largest and fastest-growing layer of the AI stack (inference), compounding ~20x with elite customers and top-tier capital, positioned to be the neutral, multi-cloud serving standard as inference eclipses training in AI spend.
- Inference is where AI value and volume are shifting; Baseten sits directly in that flow with >1B calls/day
- ~$200M -> ~$600M ARR in a single quarter-run-rate step is rare, demand-driven, real usage
- Neutral multi-cloud stance (18 clouds) makes it the Switzerland of serving vs. lock-in hyperscalers
- Marquee reference customers (Cursor, Notion, Abridge, Clay) validate production-grade reliability
- NVIDIA as both investor ($150M Series E) and supplier signals privileged access to scarce hardware and roadmap
- Software/DX + Truss create a genuine developer moat and low-friction adoption funnel
- Deep war chest (~$2.085B raised, incl. a $1.5B Series F) funds GPU pre-commitments and price aggression rivals can't all match
Baseten is a low-margin GPU reseller in a crowded, deflationary market, priced at ~22x ARR with undisclosed economics, exposed to hyperscaler bundling, NVIDIA dependence, customer in-sourcing, and relentless cost-per-token deflation.
- Gross margins are undisclosed and likely ~50% or worse due to GPU COGS - not a durable software-margin business
- At least four well-funded rivals (Together at ~$1.15B bookings/$8.3B, Fireworks, Modal, Replicate) plus Groq attack the same spend, driving price wars
- Hyperscalers can bundle inference near cost to defend far larger cloud contracts
- Biggest, most sophisticated customers are exactly the ones likely to in-source serving at scale
- ~$13B on ~$600M ARR bakes in flawless execution; any growth deceleration re-rates hard
- Dependence on NVIDIA silicon and rented capacity means little control over its own cost base
- Cost/token deflation means Baseten must keep cutting prices just to hold share - a treadmill
What it is worth
ARR multiple vs. private AI-infra comps (no public price; last-round-priced).
Thin/undisclosed margins + relentless cost/token deflation + hyperscaler bundling drive a re-rate; at ~22x ARR a growth stumble could halve the implied valuation in a down AI-infra tape. Not financial advice.
~$11B-$13B is a fair last-round clearing price for the growth, but demands continued hyper-growth; modest deceleration or a price war compresses the multiple even if ARR still rises.
Sustained multi-x growth toward $1B+ ARR with expanding gross margin (serving-stack + newer silicon) supports the $13B mark and a higher IPO-era valuation; inference becomes a durable, defensible category and Baseten a top-2 independent.
Series F priced at up to $13B post-money ($1.5B raised across two tranches, $11B and $13B) on ~$600M ARR = ~18x-22x ARR. Rich but not an outlier among fast-growing inference peers - Fireworks is reportedly in talks near ~$15B on only ~$315M ARR (~48x), while Modal is ~$2.5B on ~$50M ARR and Together is $8.3B on >$1.15B bookings (~7x). The multiple is underwritten by ~20x growth and category leadership, not by proven margins - gross/operating economics remain undisclosed.
SWOT
Strengths
- Category-leading growth (~20x YoY) on genuine usage, not just funding
- Best-in-class developer experience and Truss open-source framework as a top-of-funnel
- Multi-cloud, multi-GPU sourcing reduces single-vendor capacity risk and enables cost arbitrage
- Strong strategic backing incl. NVIDIA ($150M in the Series E) plus deep-pocketed crossover investors (Altimeter, Conviction, Spark, Sands, Wellington, Durable, D.E. Shaw, IVP, Greylock)
- Enterprise-grade compliance (HIPAA, SOC 2) unlocking regulated verticals
Weaknesses
- Undisclosed and structurally thin gross margins (GPU COGS); not a 70%+ SaaS profile
- Heavy dependence on NVIDIA silicon and on rented capacity it does not own
- Customer concentration risk among a handful of hyper-scaling AI-native accounts that could in-source or move to hyperscalers
- No moat in raw compute; differentiation is software/DX which incumbents can replicate
- Valuation (~22x ARR) leaves no room for a growth stumble
Opportunities
- Inference TAM expanding as agentic/multi-model workloads multiply per-request compute
- Open-weight model proliferation pushes more workloads to neutral serving platforms
- Move up-stack into training, fine-tuning, evals, and agent orchestration to raise ACV and stickiness
- Regulated-industry expansion (healthcare, finance) where compliance + private deployment matter
- Sovereign / on-prem / VPC deployments as enterprises demand data residency
Threats
- Hyperscalers (AWS Bedrock/SageMaker, Google Vertex, Azure AI) bundling inference at cost to lock in cloud spend
- Well-funded direct rivals (Together AI, Fireworks AI, Modal, Replicate, Groq) competing on price and speed
- Cost/token deflation compressing the spread and forcing perpetual price cuts
- GPU supply/pricing shocks and NVIDIA allocation dynamics
- Model labs offering first-party inference APIs that bypass third-party serving entirely
- AI-infra funding/valuation correction if inference demand or margins disappoint
Moats, dependencies & bottlenecks
Moats
Sticky at the developer/workflow layer, but DX advantages are replicable by well-resourced incumbents.
Sourcing and orchestration scale create availability + cost-arbitrage advantages, but not a hard technical lock-in.
Once mission-critical latency-sensitive workloads (esp. HIPAA/regulated) run on Baseten, migration is costly.
Speculative decoding, cold-start and throughput wins are real but continuously eroded as the whole field optimizes on the same silicon.
~$2.085B raised plus NVIDIA backing aids GPU access; a scale/access edge more than a structural moat.
Dependencies
Supplier + investor GPU hardware supply, allocation, pricing, and roadmap; NVIDIA also invested ~$150M in the Jan 2026 Series E.
CoreWeave, Lambda, Crusoe) Compute supplier / landlord Rents the GPU capacity it resells; these vendors are simultaneously competitors.
Product input / demand driver Serving demand depends on a healthy supply of capable open models to run.
A few scaling accounts likely drive an outsized revenue share; in-sourcing or churn would sting.
GPU pre-commitments and land-grab pricing rely on continued access to cheap growth capital and possibly GPU-backed financing.
Advantages
- First-mover brand in developer-friendly production inference
- Neutral multi-cloud positioning vs. lock-in hyperscalers
- Elite reference customer roster
- Privileged NVIDIA relationship and large capital base
- Proven ability to scale to >1B calls/day reliably
Weaknesses
- No ownership of the underlying compute or silicon
- Undisclosed, structurally thin margins
- Highly contested market with price-war dynamics
- Valuation leaves no margin for error
- Exposure to a handful of large customers that can in-source
Bottlenecks
- GPU availability and cost - the binding constraint on both growth and margin
- Cost-per-token deflation forcing continuous price cuts
- Ability to keep gross margin above zero-spread as competition intensifies
- Talent for low-level inference/systems optimization
- Customer concentration - scaling with (and retaining) a small set of large accounts
Top signals & trends
Top signals
Investor conviction that inference is a durable, expanding category and Baseten a leader.
Strategic supplier endorsement + hardware-access signal.
Demand is real and accelerating, not funding-manufactured.
Staggered investor commitments / some caution on top-end pricing amid a frothy AI-infra market.
Crowded, well-capitalized field points to intensifying price competition.
Deflation pressures the spread that token-priced revenue depends on.
Trends
Expands Baseten's core TAM.
Drives demand for neutral third-party serving.
Grows usage/TAM but compresses per-unit pricing and spread.
Competitive pressure from customers' existing cloud vendors.
More inference calls per task lifts volume.
Improves throughput/cost, aiding margin defense - but available to all rivals too.
Ecosystem & competitor graph
Suppliers feed the company; customers pull from it. Line thickness shows the strength of each tie (supply-chain dependency, customer earnings contribution). Hover to isolate a tie.
Primary GPU silicon supplier and Series E investor (~$150M) - the single most critical dependency.
GPU neocloud capacity provider (and partial competitor).
Cloud GPU capacity across the multi-cloud fleet.
Cloud GPU/TPU capacity supplier.
Cloud GPU capacity supplier.
Independent GPU cloud providers contributing to the 18-cloud sourcing base.
AI code editor; high-volume inference customer.
AI features served at scale.
Healthcare AI scribing; HIPAA-regulated inference workload.
GTM/data-enrichment AI product running inference on Baseten.
Direct rival; token-selling inference/training cloud. Raised $800M Series C at $8.3B (Jul 2026); reported >$1.15B annualized bookings - the scale leader among independents.
Direct rival; fast token-serving platform. ~$315M annualized revenue (Feb 2026, +416% YoY); in talks to raise at ~$15B valuation (a very rich ~48x).
Closest model-of-business analog (sells deployments, not just tokens); ~$50M ARR, reportedly in talks to raise at ~$2.5B.
Developer-friendly model hosting/API marketplace; overlaps on ease-of-deployment.
Custom LPU inference hardware + cloud; competes on ultra-low latency/cost-per-token.
Low-cost open-model serving; price competitor on token APIs.
SageMaker + Bedrock managed inference; can bundle at cost to defend cloud spend.
Vertex AI managed inference + TPUs; hyperscaler bundling threat.
Azure AI / Foundry managed inference tied to OpenAI ecosystem.
Public GPU 'neocloud'; both a capacity supplier and a competitor moving up-stack into managed inference.