
LlamaIndex
Open-source core (Python/TypeScript framework, permissive MIT-style license) monetized via a managed cloud platform: usage/credit-based document parsing (LlamaParse) plus subscription tiers for LlamaCloud (managed knowledge/retrieval), and enterprise/self-hosted contracts. Classic open-source-to-commercial (OSS funnel -> managed-service conversion).
Valuations are third-party estimates (CB Insights/PitchBook); LlamaIndex has not disclosed official post-money figures. The May 2025 Databricks/KPMG strategic minority investments were undisclosed in amount and did not reprice the company publicly, so the trail holds the ~$93M Series A mark through that date.
Earnings, margins, COGS & capex
Private, seed/Series-A-stage OSS-infra company. Total disclosed equity raised ~$27.5M ($8.5M seed Jun 2023 led by Greylock, at ~$45M post-money; $19M Series A Mar 2025 led by Norwest Venture Partners with Greylock) plus undisclosed May 2025 strategic minority investments from Databricks Ventures and KPMG Ventures. Post-money ~$93M at Series A (third-party estimate) - roughly a 2x (~107%) step-up on the ~$45M seed post-money. No audited revenue or margins are public; a third-party estimate puts 2025 revenue near $10.9M. Traction is demonstrated through usage metrics rather than disclosed financials: 500M+ documents processed, ~4M monthly package downloads, ~200,000 LlamaCloud users, and a waitlist of 10,000+ organizations including ~90 Fortune 500 companies.
Revenue trend
Margins
Compute/LLM cost of LlamaParse is a structural drag vs pure SaaS
Growth-stage burn
VC-funded
COGS structure
Primary COGS is third-party LLM inference and compute for the document-parsing/extraction pipeline (LlamaParse/LlamaExtract) plus cloud hosting for LlamaCloud. Per-page parsing is priced in credits (~1,000 credits = $1.25; roughly $0.00125-$0.05625 per page by document complexity), so COGS scales with document volume and model choice - a genuine gross-margin exposure vs seat-based SaaS peers.
Capex
Not disclosed; minimal - cloud-native, no owned infrastructure. Investment is in R&D headcount, not physical assets.
Latest earnings
Not applicable.
No public guidance.
- Documents processed
- 500M+ (cumulative, 2025)
- Monthly package downloads
- ~4M
- LlamaCloud users
- ~200,000
- Waitlist orgs
- 10,000+ (incl. ~90 Fortune 500)
- Total disclosed raised
- ~$27.5M + undisclosed strategic (thru 2025)
- Headcount
- 30+ (mid-2024; later figures not disclosed)
Growth drivers
- Enterprise RAG / knowledge-agent adoption - Fortune 500 building internal AI assistants over private documents
- LlamaParse as a best-in-class parser for messy, complex enterprise PDFs (tables, forms, scans) - a wedge product with clear willingness-to-pay
- LlamaCloud managed platform converting free OSS users into paying managed-service customers
- Distribution via Databricks (Data Intelligence Platform) and KPMG (services/deployment channel) strategic partnerships
- Very large top-of-funnel: ~4M monthly downloads and ~200k LlamaCloud users feeding paid conversion
Bull & bear
LlamaIndex owns the developer default for building over unstructured enterprise data, and is converting that mindshare into a real managed business (LlamaCloud/LlamaParse) with strategic distribution through Databricks and KPMG - a credible path to being the 'context/retrieval layer' of the enterprise AI stack.
- Genuine category leadership and OSS distribution moat - ~4M monthly downloads and 500M+ documents processed is a top-of-funnel most infra startups never achieve
- LlamaParse solves a concrete, painful, paid problem (messy enterprise document ingestion) that pure-LLM APIs still do poorly - a durable wedge with clear willingness-to-pay
- Strategic investors Databricks and KPMG bring enterprise distribution and credibility, not just capital, into exactly the regulated verticals with the biggest RAG budgets
- Enterprise agentic-AI adoption is early; 10,000+ waitlisted orgs incl. ~90 Fortune 500 suggests demand well ahead of monetized revenue
- Lean cap table (~$27.5M raised) means a modest ARR ramp and any strategic acquisition (Databricks, Snowflake, a hyperscaler, ServiceNow) could deliver a strong multiple on invested capital
A thinly-capitalized ~$10.9M-revenue framework company sitting directly in the path of hyperscalers and model providers who are commoditizing RAG and native retrieval - with its own CEO conceding the 'framework era' is ending. Monetization at scale is unproven and gross margins are structurally exposed to LLM compute costs.
- The core framework is being commoditized from above (native long-context, model-provider retrieval/tools) and below (hyperscaler RAG bundles) - the reason to depend on a third-party framework is shrinking
- Revenue (~$10.9M est.) is small and undisclosed/unaudited; OSS-to-paid conversion at scale is not yet demonstrated
- Databricks is simultaneously an investor and a competitor (Mosaic AI agent framework) - the strategic tie could cap independence or presage an acqui-exit rather than an independent breakout
- Gross margins are pressured by per-page LLM/compute COGS, unlike seat-based SaaS peers - scaling parsing volume may not scale margin
- At a ~$93M post-money on ~$10.9M revenue (~8-9x sales) the entry is not cheap for a business facing this much platform-risk from far better-capitalized incumbents
What it is worth
Last-priced-round anchor plus revenue-multiple sanity check (no public market comps apply; private, pre-IPO).
Down-round or distressed acqui-hire if native model-provider retrieval and hyperscaler RAG bundles commoditize the framework faster than LlamaCloud monetizes - the 'end of the framework era' scenario.
Holds around the ~$93M Series A mark; incremental step-ups tied to demonstrated paid-conversion of the OSS base.
Rerating well above $93M if LlamaCloud ARR compounds and LlamaParse becomes the enterprise default for document ingestion, supported by Databricks/KPMG distribution - or a premium strategic acquisition given the lean cap table.
Last priced round: ~$93M post-money at the $19M Series A (Norwest lead, Mar 2025; third-party estimate), up from ~$45M seed post-money (Greylock, Jun 2023). On an estimated ~$10.9M 2025 revenue that implies ~8-9x sales - reasonable-to-full for a category-leading AI infra name, but rich given platform/commoditization risk. Undisclosed Databricks/KPMG strategic investments (May 2025) were not repriced publicly. No public price exists; any exit is most plausibly a strategic acquisition (Databricks, Snowflake, a hyperscaler, ServiceNow) rather than a near-term IPO. Not financial advice.
SWOT
Strengths
- Category-defining open-source brand in RAG/data-framework tooling with massive developer mindshare (~4M monthly downloads)
- LlamaParse is a differentiated, hard-to-replicate wedge for complex-document parsing with direct monetization
- Blue-chip validation — Norwest + Greylock backing, plus strategic equity from Databricks and KPMG (distribution, not just capital)
- Founder-market fit — Jerry Liu (CEO, ex-Two Sigma/Uber/Quora) and Simon Suo (CTO) built the de facto standard from the GPT Index OSS project
Weaknesses
- Small scale and undisclosed, likely-negative financials — ~$10.9M estimated revenue is tiny vs the enterprise-AI TAM being claimed
- OSS-to-commercial conversion is unproven at scale - most usage is free
- COGS exposure to third-party LLM/compute pricing pressures gross margin vs seat-based SaaS
- Thin capitalization (~$27.5M disclosed) relative to well-funded rivals (LangChain, Databricks, hyperscalers)
Opportunities
- Enterprise agent/knowledge-management spend is early and expanding; document-heavy verticals (finance, legal, insurance, healthcare) are natural fits
- Upsell OSS base into LlamaCloud managed and enterprise/self-hosted contracts
- Deepen Databricks/KPMG channels into large regulated-enterprise deployments
- Expand from retrieval into full agentic workflows / extraction (LlamaExtract) as the 'context layer' for enterprise AI
Threats
- Commoditization — hyperscalers (AWS Bedrock Knowledge Bases + Textract, Azure AI Search, Google Vertex AI) bundle RAG/parsing into their clouds for near-zero marginal cost
- 'End of the framework era' risk — the CEO has publicly argued frameworks are being subsumed as agent loops and long-context models mature; model providers ship native retrieval/tools/long-context that erode the framework's reason to exist
- Direct competition from LangChain/LangGraph (broader agent orchestration) and Databricks' own Mosaic AI (an investor that is also a competitor)
- Vector-DB and doc-AI point players (Pinecone, Weaviate, Unstructured.io) competing for the same infra budget
Moats, dependencies & bottlenecks
Moats
~4M monthly downloads and category-defining brand; but developer defaults can shift quickly as model capabilities absorb framework functionality.
Best-in-class on complex/messy docs today; a real technical lead, but a replicable one as OCR+LLM parsing improves industry-wide.
Distribution + credibility into regulated enterprise; double-edged given Databricks also competes.
Once embedded in production RAG pipelines there is stickiness, but the framework layer is deliberately swappable.
Dependencies
Anthropic, and open-weight models) Supplier / platform Both a COGS input and a strategic threat - the same providers ship native retrieval/long-context that can disintermediate the framework.
LlamaCloud runs on hyperscaler compute - the hyperscalers are also competitors bundling rival RAG services.
The OSS funnel is the growth engine; community goodwill and momentum must be sustained against LangChain and native tooling.
Investor / partner / competitor Strategic investor and distribution channel that also builds a competing agent/retrieval stack (Mosaic AI).
Revenue is levered to enterprises continuing to build (vs buy) internal AI assistants.
Advantages
- First-mover / category-owner brand in RAG data frameworks
- Concrete paid wedge (LlamaParse) beyond a free framework
- Strategic-investor distribution (Databricks, KPMG) into regulated enterprise
- Founder-market fit and deep RAG/agent technical credibility
Weaknesses
- Tiny, undisclosed, likely-unprofitable financials relative to the TAM narrative
- High platform-risk from LLM providers and hyperscalers commoditizing the layer
- Thin capitalization vs competitors
- Compute-cost exposure in COGS undermines classic SaaS margin structure
Bottlenecks
- Monetizing a predominantly free OSS user base - converting downloads to paid LlamaCloud/enterprise seats
- Gross-margin control as document-parsing volume (and LLM compute cost) scales
- Enterprise sales motion and support capacity at ~30-person scale vs incumbents' large field orgs
- Defending the framework's relevance as model providers ship native retrieval/agent primitives
Top signals & trends
Top signals
Validates enterprise relevance and opens distribution; terms undisclosed.
Priced up from ~$45M seed post-money (Jun 2023) to ~$93M post-money on milestone execution.
Honest but signals the core framework's defensibility is under question; pivot toward the 'context layer' framing.
Large funnel is bullish; the conversion gap is the open question.
Trends
Expands the paid TAM for retrieval/parsing infrastructure.
Erodes the need for a separate framework/retrieval layer - the central bear thesis.
Commoditizes the managed offering at near-zero marginal cost to the cloud.
Favors LlamaParse/LlamaExtract as specialized, higher-accuracy tools.
Ecosystem & competitor graph
Suppliers feed the company; customers pull from it. Line thickness shows the strength of each tie (supply-chain dependency, customer earnings contribution). Hover to isolate a tie.
LLM/embedding inference - a core COGS input and a disintermediation risk. Private.
LLM provider integrated across the framework. Private.
Underlying GPU compute powering all inference in the pipeline (via clouds).
Hosting/compute for LlamaCloud - simultaneously suppliers and competitors.
Enterprises building internal knowledge agents 10,000+ waitlisted orgs incl. ~90 Fortune 500; concentrated in document-heavy verticals (finance, legal, insurance, healthcare, professional services).
The OSS free-tier funnel that seeds paid LlamaCloud/LlamaParse conversion.
Strategic investor deploying LlamaIndex in client-facing enterprise AI engagements.
Closest direct rival - broader agent-orchestration framework with its own commercial platform (LangSmith); competes for the same developer default. Private.
Strategic investor that also ships a competing agent/retrieval stack on its Data Intelligence Platform. Private (frequent IPO candidate).
Azure AI Search + Foundry/Copilot stack bundles enterprise RAG natively into the dominant enterprise cloud.
Bedrock Knowledge Bases + Textract provide managed RAG and document parsing directly competing with LlamaCloud/LlamaParse.
Vertex AI Search / Document AI + long-context Gemini bundle retrieval and parsing.
Managed vector database; overlaps on the retrieval-infra budget. Private.
Document ingestion/ETL for LLMs - directly competes with LlamaParse on parsing. Private.
Cortex Search/AI brings RAG to data already sitting in the warehouse.
Vector search inside the operational database - reduces the need for a separate retrieval layer.
Elasticsearch vector/hybrid search is an established enterprise retrieval substrate.
Open-source and managed RAG/vector alternatives competing for the same developer and enterprise budget. Private.