
Groq
AI inference accelerators (LPU) + inference neocloud (GroqCloud); fabless chip + pay-per-token cloud + hardware/licensing
Priced rounds C/D/E are firm; Series C is a disclosed '>$1B' unicorn threshold (exact post-money not given). The Dec-2025 Nvidia ~$20B figure is a licensing+acqui-hire deal value (patents/software/team), not a priced equity round — tagged secondary. Seed (2017) and the Jun-2026 $650M round had undisclosed valuations and are omitted.
Earnings, margins, COGS & capex
Groq does not publish audited financials. Known operating signals (vintage-stamped): 5M+ developers and thousands of AI companies on GroqCloud (Jun 2026); trillions of tokens processed weekly; ~75% of Fortune 100 reportedly hold platform accounts (Sacra). Revenue mix is pay-per-token GroqCloud + GroqRack hardware + the new (non-recurring) Nvidia licensing cash. The ~$17B Nvidia payment stream (Dec 2025) is a one-time IP monetization, not a run-rate; do not capitalize it as recurring revenue. Capex is heavy and external-funded: Saudi/HUMAIN $1.5B commitment (Feb 2025) plus the $650M Jun 2026 raise fund the 200MW build-out. Gross/operating/FCF margins, COGS, and net cash are all undisclosed.
Income statement — where each revenue dollar goes
% of revenueOf every $1 of revenue, ~83¢ is cost of goods and ~17¢ operating expense, leaving ~0¢ of operating profit.
Revenue trend
Margins
COGS structure
Not disclosed (private company).
Capex
Not disclosed.
Growth drivers
- Inference is the largest and fastest-growing AI-compute segment — Groq is a recognized brand with 5M+ developers and trillions of tokens/week of real usage.
- The Nvidia deal de-risked the balance sheet (~$17B cash through 2026) and validated the IP at ~$20B — the standalone equity now trades well below that mark.
- Sovereign/strategic demand (HUMAIN $1.5B, Aramco, Bell Canada) underwrites multi-year capacity off-take that most neocloud peers lack.
- Day-zero open-model support (Llama 4) and pay-per-token pricing give a defensible developer-acquisition flywheel as models proliferate.
Bull & bear
A re-capitalized, focused inference neocloud riding the inference mix-shift with a huge developer funnel, sovereign demand, and a $17B Nvidia cash cushion — cheap optionality if the cloud business compounds.
- Inference is the largest and fastest-growing AI-compute segment; Groq is a recognized brand with 5M+ developers and trillions of tokens/week of real usage.
- The Nvidia deal de-risked the balance sheet (~$17B cash through 2026) and validated the IP at ~$20B — the standalone equity now trades well below that mark.
- Sovereign/strategic demand (HUMAIN $1.5B, Aramco, Bell Canada) underwrites multi-year capacity off-take that most neocloud peers lack.
- Day-zero open-model support (Llama 4) and pay-per-token pricing give a defensible developer-acquisition flywheel as models proliferate.
Groq sold its crown jewels and its team to its biggest rival; what's left is a capital-intensive, low-margin neocloud competing against Nvidia, hyperscaler ASICs, and a freshly-public Cerebras, with no audited profitability and an IPO path that the Nvidia deal arguably foreclosed.
- The people who built the LPU moat — Ross, Madra, and team — now work at Nvidia, which holds a non-exclusive license to the same IP and is embedding it in Rubin systems.
- Revenue is unaudited, modest (~$500M 2025E), and the heroic figure (the ~$17B Nvidia cash) is non-recurring — strip it out and the operating business is unproven.
- Inference token pricing is deflationary; competing on cost against Nvidia's installed base and hyperscaler in-house silicon is a margin trap.
- Down-round risk and IPO foreclosure: the deal showed a well-funded, customer-rich startup still couldn't reach escape velocity in Nvidia's shadow.
What it is worth
Private — no public price. Anchored to last priced round ($6.9B Sept-2025 Series E), the Nvidia ~$20B IP-license mark, secondary marks ($3.6B–$6.9B), and a revenue-multiple sanity check on unaudited ~$500M 2025E (~7x, per Sacra). Reverse-DCF logic: the current standalone mark prices the cloud business well below the $20B IP value, implying the market discounts the post-acqui-hire operating entity heavily.
~$2–4B
down round confirmed; margin compression and Nvidia/hyperscaler competition shrink the standalone franchise toward secondary-mark lows.
~$5–7B
holds near the last disclosed $6.9B mark as a well-funded but unproven standalone inference neocloud; flat-to-modest re-rate.
~$10B+
GroqCloud compounds paid token revenue, sovereign off-take ramps, $17B Nvidia cash funds 200MW; re-rates toward Cerebras-style public multiples.
SWOT
Strengths
- Differentiated deterministic LPU architecture (SRAM-on-chip, compiler-scheduled, no caches/branch prediction) delivering class-leading low-latency token throughput — validated by Nvidia paying ~$20B to license it.
- Large developer funnel — 5M+ developers, trillions of tokens/week, day-zero model support (Meta Llama 4 via the official Llama API).
- Deep-pocketed sovereign/strategic backers — Saudi Arabia/HUMAIN ($1.5B), Aramco Digital data-center partnership, Bell Canada sovereign-AI deployment, plus a $17B Nvidia cash backstop.
- Pay-per-token neocloud model is the recurring-revenue shape investors prize in 2026; capacity scaling to 200MW by 2027.
Weaknesses
- Founder/CEO Jonathan Ross (TPU architect) and president Sunny Madra left for Nvidia in Dec 2025 — the core inference IP team is gone; the company is re-staffing leadership (new CEO Simon Edwards; new COO/CTO/CPO hires in 2026).
- Non-audited, modest operating revenue (~$500M 2025E) against a heavy-capex build-out — the business runs on external capital, not self-funding.
- No public margin disclosure — inference token pricing is in a race to the bottom against Nvidia, hyperscaler silicon, and other neoclouds.
- Strategic dependency on third-party open models (Llama) and on a foundry roadmap (Samsung Taylor, TX, 4nm) it does not control.
Opportunities
- Inference is already ~2/3 of AI compute, projected ~80% by 2027 — the fastest-growing layer; a focused inference-cloud pure-play can ride that mix shift.
- Sovereign-AI demand (Gulf, Canada, EU) for non-US-hyperscaler, low-cost inference capacity plays to Groq's data-center partnerships.
- Re-use the $17B Nvidia cash to fund the 200MW build-out and buy share before competitors scale.
- Non-exclusive license means Groq can still license LPU IP to other partners and keep selling GroqRack.
Threats
- Nvidia — now both Groq's largest counterparty AND its dominant competitor (80–90% AI-silicon share) — is integrating Groq's LPU IP into Rubin/AI-factory systems, eroding Groq's latency edge.
- Cerebras (public since May 2026, ~$56B) and a wave of inference startups (Etched, d-Matrix, Fireworks, Baseten) plus hyperscaler ASICs (Google TPU, AWS Trainium/Inferentia, Microsoft Maia) compete for the same token spend.
- Customer concentration / sovereign dependency — a large share of committed capacity is tied to Saudi/HUMAIN and a handful of strategic deals.
- Token-price deflation and power/data-center cost inflation can crush inference-cloud gross margins.
Moats, dependencies & bottlenecks
Moats
SRAM-resident, compiler-scheduled) — a genuinely different inference design, the asset Nvidia paid ~$20B to license.
day-zero model support).
Aramco Digital, Bell Canada) that are hard for newcomers to replicate.
Software/compiler stack tuned to the LPU's deterministic execution model.
Dependencies
Samsung (Taylor, Texas, 4nm) for next-gen LPU manufacturing.
especially Meta Llama, for GroqCloud demand.
Nvidia cash) to fund the 200MW build-out.
HBM/SRAM and packaging supply chain.
Advantages
- Latency/throughput leadership per chip on transformer inference.
- Lower cost-per-token positioning (e.g., Saudi inference cost claims).
- Brand recognition with developers and day-zero model launches.
- Balance-sheet strength post-Nvidia relative to fire-sale peers (SambaNova ~$1.6B).
Weaknesses
- Loss of founding architecture team and CEO to a direct competitor.
- No audited financials or disclosed margins; unproven standalone profitability.
- Heavy reliance on a few sovereign/strategic relationships and on third-party models.
- Competing head-on with its own largest counterparty (Nvidia).
Bottlenecks
- Capacity/power: scaling to 200MW by 2027 is capital- and grid-constrained.
- Talent: rebuilding deep silicon/architecture bench after the Nvidia acqui-hire.
- Foundry allocation at Samsung 4nm against larger buyers.
- Token gross margin under price competition limits self-funded growth.
Top signals & trends
Top signals
Valuation undisclosed; secondary marks diverge as low as ~$3.6B (Sacra), implying possible down-round post-Nvidia deal.
Funds the 200MW build-out and de-risks the balance sheet if installments land on schedule.
New team from SpaceX (xAI)/Meta/Microsoft must rebuild the silicon roadmap without the founding architects.
5M+ developers is reach; the question is durable, profitable, paid run-rate vs free usage.
Erodes Groq's standalone latency differentiation and competes for the same inference demand.
Trends
~2/3 of AI compute is inference in 2026, ~80% by 2027 — Groq is a pure-play on this shift.
Compresses inference-cloud gross margins; favors the lowest-cost/largest-scale operators.
Demand for non-hyperscaler, low-cost inference capacity underwrites Groq off-take.
Google TPU, AWS Trainium/Inferentia, Microsoft Maia, and Nvidia Rubin squeeze independents.
3–5 more deals expected by end-2026; validates the category but signals standalone difficulty.
Ecosystem & competitor graph
Suppliers feed the company; customers pull from it. Line thickness shows the strength of each tie (supply-chain dependency, customer earnings contribution). Hover to isolate a tie.
Foundry for next-gen LPU (Taylor, TX; 4nm).
Data-center partner for the Dammam/EMEA inference hub (non-US; analysis only).
Representative data-center networking supplier class for large inference build-outs.
Representative power/thermal/data-center infrastructure supplier for 200MW scaling.
Official Llama API runs inference through Groq LPUs (Llama 4).
$1.5B sovereign-AI commitment; anchor capacity off-taker (non-US, private).
Sovereign-AI deployment across multiple Canadian sites for gov/enterprise.
Inference-provider integration routing developer demand to GroqCloud (private).
Dominant AI silicon (80–90% share); now licenses Groq's LPU IP and employs its founders — counterparty and competitor.
Wafer-scale inference; IPO'd May 2026 at ~$56B FDV — the public comp and direct inference-speed rival.
Instinct MI-series GPUs targeting inference; the credible #2 merchant accelerator.
TPU v-series — in-house inference silicon and a cloud that competes for the same token spend.
AWS Inferentia/Trainium — hyperscaler ASICs undercutting third-party inference clouds.
Maia accelerators + Azure AI — in-house inference silicon and a competing cloud.