
Fireworks AI
Usage-based, developer-first API. Revenue across serverless per-token inference, fine-tuning (per training token), reinforcement fine-tuning (per GPU-hour), and dedicated GPU deployments (per GPU-second/hour). Bottoms-up developer entry, top-down enterprise expansion.
Chronological priced primary rounds. The Series A ~$0.1B post-money is an estimate (round size $25M disclosed; post-money not officially published). The 2026-05 ~$15B point is a reported, still-open round (not closed as of the Jul 2026 asOf) and terms remain subject to change; a stray 'nears $100B' headline (Tiger Brokers) is an unsupported outlier and is excluded. Implied step-up from Series C to the rumored round is ~3.75x in ~7 months.
Earnings, margins, COGS & capex
Private inference platform scaling revenue faster than almost any peer: from a ~$280M annualized run-rate at the Oct-2025 Series C (~$305M by end-2025) to a reported ~$800M by May 2026 (Sacra estimate), on 10,000+ customers and daily token volume that grew from ~10 trillion (Oct 2025) to ~30 trillion (mid-2026). Economics are GPU-cost-dominated: ~50% gross margin today, with margin expansion the central bet. No audited public statements exist; all figures are round/press-disclosed or third-party estimates and should be treated as management/estimate figures.
Income statement — where each revenue dollar goes
% of revenueOf every $1 of revenue, ~50¢ is cost of goods and ~50¢ operating expense, leaving ~0¢ of operating profit.
Revenue trend
Margins
improving toward stated ~60% target
unknown; heavy GPU + R&D spend implies likely unprofitable / reinvesting
skewed — a few AI-native heavy users (e.g. Cursor) dominate volume
COGS structure
GPU compute is the dominant cost: NVIDIA hardware procurement (owned + rented from neoclouds/hyperscalers), datacenter capacity planning, and regional infrastructure. This is what caps gross margin near 50% and makes utilization/hardware-efficiency the key lever.
Capex
Not separately disclosed. Model is a mix of owned GPU fleet and rented capacity; capital intensity tracks GPU acquisition and reserved-capacity commitments. A shift toward more owned/committed capacity raises capex but can improve unit economics at scale.
Latest earnings
n/a
No formal guidance. Directional signal: reported ~$800M annualized run-rate (May 2026) and an in-progress raise at ~$15B implies investors underwriting continued hyper-growth.
- Annualized revenue (May 2026, Sacra)
- ~$800M
- Tokens processed / day (mid-2026)
- ~30 trillion
- Customers
- 10,000+
- Gross margin
- ~50% (targeting ~60%)
- Total equity raised
- $327M+ through Series C
Growth drivers
- Secular explosion in AI inference token volume (agentic apps, coding assistants, RAG) — Fireworks' daily tokens grew ~3x to ~30T in under a year
- Open-weight model wave (DeepSeek, Kimi/Moonshot, Qwen, Llama, Mistral) that customers want served cheaply and fast — Fireworks' core wedge
- AI-native flagship customers scaling usage — Cursor, Notion, Sourcegraph, Uber, DoorDash, Shopify, Upwork, Samsung
- Expansion up-market from serverless into dedicated deployments, fine-tuning, and reinforcement fine-tuning (higher-value, stickier)
- Proprietary inference-optimization stack (custom kernels, FireAttention, speculative decoding, disaggregated serving) lowering cost-per-token
Bull & bear
Fireworks is the leading independent, model-neutral inference platform at the exact moment inference (not training) becomes the dominant AI compute spend — with a world-class systems team, hyper-growth revenue, and marquee AI-native customers that make it the default serving layer for open-weight models.
- Inference is the durable, recurring AI spend layer; Fireworks' ~3x token growth in a year shows it is compounding with the market
- PyTorch-pedigree team + proprietary optimization stack = a real, defensible cost/latency edge, not a reseller
- Model-neutral positioning wins as enterprises hedge against any single lab and adopt cheaper open-weight models
- Land-and-expand into fine-tuning, RFT, and dedicated deployments raises margin and switching costs over time
- Reported ~$800M run-rate at ~50% gross margin is a genuinely large, fast, real-revenue business — not a pre-revenue story
- A ~$15B raise (if closed) funds owned-capacity and hardware-efficiency investment that structurally lifts margins
Inference is commoditizing fast: per-token prices are nearly identical across Fireworks, Together, and Baseten, gross margins are GPU-capped near 50%, and the biggest, best-funded players (hyperscalers, NVIDIA, model labs) all want this exact market — leaving an independent middleware layer squeezed on both price and supply.
- Prices are effectively at parity with Together AI and Baseten — no pricing power, margins bleed as competition intensifies
- ~50% gross margin means limited operating leverage; scaling revenue also scales GPU cost near-linearly
- Hyperscalers can bundle inference into existing enterprise cloud spend and undercut a standalone provider
- Large customers (Cursor et al.) have every incentive to in-source serving with vLLM/SGLang once volume justifies it — concentration cuts both ways
- Dependent on NVIDIA supply and third-party capacity, with little control over the most important input cost
- A ~3.75x valuation step in 7 months to ~$15B prices in flawless execution; any AI-capex cooling repricing hits hard, and the round is not yet closed
What it is worth
Private-market, last-round + revenue-multiple triangulation (no public price exists).
Down-round risk toward the Series C ~$4B or below if AI-capex sentiment cools, hyperscalers/NVIDIA compress inference pricing, or large customers in-source serving — the middleware layer gets squeezed on price and supply.
~$4-8B
real, fast-growing revenue but commoditization and ~50% margins argue for a multiple below the peak AI-infra hype; value hinges on holding the volume lead over Together and moving up-stack.
~$15B+
if the raise closes and revenue keeps compounding toward a $1B+ run-rate with margins trending to 60% — inference-market leadership justifies a scarcity premium.
Series C set ~$4B post-money (Oct 2025) on a ~$280M run-rate (~14x forward revenue). A reported ~$15B target (Bloomberg, May 2026) against a ~$800M run-rate is ~19x — a premium multiple that prices in continued hyper-growth and margin expansion, and the round is not yet closed. Because gross margin is GPU-capped near 50%, this is valued much closer to a high-growth software business than the capital-intensive economics arguably warrant. Private/pre-IPO — not investable as a public equity; figures are round-disclosed or third-party estimates (Sacra), not audited. Not financial advice.
SWOT
Strengths
- Elite founding/engineering team from Meta's PyTorch org (CEO Lin Qiao co-created PyTorch/Caffe2) — deep credibility and talent in model-serving systems
- Proprietary inference-optimization stack delivering speed/cost leadership on open-weight models
- Extraordinary revenue growth (~$280M to ~$800M run-rate in ~7 months) with marquee AI-native customers
- Broad, timely open-model catalog (Llama, DeepSeek, Kimi, Qwen, Mistral) served day-one — the neutral 'Switzerland' for open models
- Full-stack product spanning serverless, fine-tuning, RFT, and dedicated deployments enables land-and-expand rather than a single-SKU passthrough
Weaknesses
- Structurally GPU-cost-heavy — ~50% gross margin caps operating leverage vs software-native peers
- Heavy dependence on NVIDIA hardware and third-party cloud/neocloud capacity for supply
- Commoditization pressure — per-token prices are near-identical to Together AI/Baseten, so pricing is a race to the bottom
- Customer concentration risk — a handful of AI-native accounts likely drive a large share of tokens/revenue
- No audited financials — investors and enterprise buyers underwrite growth on self-reported / third-party estimate figures
Opportunities
- Move up-stack into higher-margin fine-tuning, RFT, agent tooling, and dedicated/VPC enterprise deployments
- Enterprise/regulated-industry expansion where data isolation and reliability command premium pricing
- Efficiency gains on Blackwell-class hardware and better utilization to reach the 60% gross-margin target
- Own more of the compound-AI/agent workflow (routing, eval, memory) to become a platform, not a passthrough API
- International/regional deployment footprint to win data-residency-constrained enterprise and sovereign workloads
Threats
- Hyperscalers (AWS Bedrock, Google Vertex, Azure) bundling inference into existing cloud relationships and undercutting on price
- Model labs (OpenAI, Anthropic, and open-weight labs' own APIs) capturing inference demand directly
- Specialized silicon (Groq, Cerebras, SambaNova) and NVIDIA's own DGX Cloud/NIM competing on cost-per-token
- Open-source serving frameworks (vLLM, SGLang) letting large customers in-source inference and cut Fireworks out
- Valuation/round risk — a ~$15B target (~3.75x in 7 months) sets a high bar; an AI-capex pullback could reprice the sector
Moats, dependencies & bottlenecks
Moats
FireAttention, speculative decoding, disaggregated serving) Moderate-to-strong real today but continuously eroded by open-source (vLLM/SGLang) catching up The core technical wedge; must keep out-engineering free frameworks to justify the spread.
Deep systems credibility drives both product edge and enterprise trust.
Serverless per-token usage is easy to leave; custom checkpoints and VPC deployments are stickier.
Being first/fastest to serve new open models (Kimi, DeepSeek, Qwen) drives developer default status.
Higher aggregate volume improves GPU utilization and unit cost — a flywheel if it holds the volume lead over Together.
Dependencies
GPUs are the dominant COGS and the binding capacity constraint; NVIDIA also competes via DGX Cloud/NIM and is a strategic investor in Fireworks.
Oracle ORCL, AWS/AMZN) Compute capacity / hosting Rented GPU capacity fills fleet gaps; pricing and availability directly hit margin.
DeepSeek, Moonshot Kimi, Alibaba Qwen, Mistral) Product supply (the models it serves) Business depends on a steady pipeline of strong open models; if open models stall or lock down, the wedge narrows.
Revenue concentration A few heavy token users likely drive outsized revenue; any in-sourcing or churn is material.
GPU-capacity buildout is capital-intensive; the model relies on ongoing equity (the ~$15B round is not yet closed).
Advantages
- Best-in-class inference speed/cost on open-weight models
- Founder/engineering pedigree (PyTorch) and systems depth
- Model-neutral, broad, day-one catalog — the default independent serving layer
- Hyper-growth revenue with real, blue-chip AI-native customers
- Full-stack product (serverless, fine-tuning, RFT, dedicated) enabling land-and-expand
Weaknesses
- GPU-capped ~50% gross margin and heavy capital intensity
- No pricing power — per-token parity with Together/Baseten
- Dependence on NVIDIA and third-party capacity for the key input
- Likely customer concentration among a few heavy token users
- Exposed to hyperscaler bundling and customer in-sourcing (vLLM/SGLang)
Bottlenecks
- GPU supply and cost — the hard ceiling on both capacity growth and gross margin
- Price commoditization across near-identical per-token markets removes pricing power
- Utilization efficiency — idle GPU time directly destroys unit economics
- Talent competition for scarce inference-systems engineers against hyperscalers and labs
- Enterprise procurement/compliance (VPC, data isolation) needed to win regulated, higher-margin accounts
Top signals & trends
Top signals
Bullish (with caution) · Investor conviction in growth, but not yet closed and sets a very high execution bar.
Among the fastest revenue ramps in AI infra.
Volume compounding with the inference market.
Confirms commoditization; differentiation must come from speed, reliability, and product depth, not price.
Margin expansion is the whole thesis and remains unproven.
Trends
Structural tailwind — inference is Fireworks' entire market.
More demand for neutral, cheap serving of open models.
Drives per-customer token growth.
Biggest players targeting the same market with bundling and captive silicon.
Lowers the cost for large customers to in-source and erodes the optimization moat.
Fuels cheap capital now; raises repricing risk if sentiment turns.
Ecosystem & competitor graph
Suppliers feed the company; customers pull from it. Line thickness shows the strength of each tie (supply-chain dependency, customer earnings contribution). Hover to isolate a tie.
GPUs (Hopper/Blackwell) — the dominant cost input and capacity constraint; also a strategic investor.
Neocloud GPU capacity provider for burst/rented fleet.
OCI GPU capacity — one of the large third-party compute sources.
Cloud infrastructure/capacity (also a competitor via Bedrock).
Supplies flagship open-weight models Fireworks serves.
Moonshot/Kimi, Alibaba Qwen, Mistral) Non-US / private model suppliers whose open weights are core inventory (analytical context only, not an investment call).
AI coding assistant — high-volume token user.
Productivity software with embedded AI features.
Code-AI/search tooling.
Enterprise AI workloads (named in Series C announcement).
Enterprise AI workloads.
Enterprise AI workloads.
Commerce AI features.
Marketplace AI features.
Closest direct rival — private, per-token open-model API with a broad catalog and fine-tune hosting; >$150M annualized revenue reported (Sacra), raised $305M Series B Feb 2025. Near-identical pricing.
Private enterprise inference platform emphasizing self-hosted/hybrid VPC deployment; raised $300M Series E at ~$5B (Feb 2026) — strong in compliance-heavy accounts.
Private; developer-friendly model-hosting/API, strong in image/multimodal and long-tail models.
Private; focused on fast image/video generation inference — overlaps in media, not core LLM.
Private; custom LPU silicon competing on ultra-low-latency, low-cost token serving.
Wafer-scale AI silicon vendor also offering fast inference; filed to IPO — competes on speed.
Private; custom-chip inference/appliance vendor targeting enterprise.
Hyperscaler managed model API bundled into AWS spend — the scale threat.
Hyperscaler model-serving with TPU cost advantage and enterprise reach.
Azure + OpenAI distribution; captive enterprise inference demand.
Supplier and competitor — offering its own managed inference microservices on the hardware Fireworks depends on.