
Zilliz (Milvus)
Open-source core (Milvus, Apache 2.0, LF AI & Data Foundation graduated project) + commercial managed cloud service (Zilliz Cloud, usage-based/pay-as-you-go on AWS, GCP, Azure across 30+ regions) + enterprise support; classic open-core COSS funnel from OSS adoption to paid cloud
Zilliz has never disclosed a valuation at any round. All valuationB figures here are explicitly labeled analyst constructs anchored to the two real primary rounds (Nov 2020 $43M, Aug 2022 $60M) and the dossier's base-case scenario; they are not reported marks and should be rendered as estimates.
Earnings, margins, COGS & capex
Zilliz is a private growth-stage infrastructure software company and publishes no financial statements. Total disclosed funding is ~$113M: a $43M Series B (Nov 2020, led by Hillhouse Capital with TrustBridge Capital, Pavilion Capital, 5Y Capital, and Yunqi, raised while primarily China-based) and a $60M Series B extension (Aug 2022) led by Prosperity7 Ventures (Saudi Aramco's $1B venture fund) with Temasek's Pavilion Capital, Hillhouse, 5Y Capital, and Yunqi Capital, announced alongside the relocation of HQ to San Francisco. No priced round has been announced since Aug 2022 and no valuation was disclosed for any round. Revenue is not disclosed; the figures in circulation are third-party estimates - ~$16.2M for 2023 (GetLatka) and ~$48.8M for 2024 (UpMarket, citing Sacra-derived estimates) - which are unaudited, unverified, and mutually imply an implausibly steep jump, so treat both with caution. Monetization runs through Zilliz Cloud consumption pricing (including AWS Marketplace pay-as-you-go) layered on top of free open-source Milvus adoption (40,000+ GitHub stars as of Dec 2025, 10,000+ enterprise teams claimed running Milvus in production).
Revenue trend
Margins
managed-cloud COSS peers typically run 60-75% at scale; Zilliz's RaBitQ compression and tiered-storage cost work suggests active COGS engineering
presumed negative; growth-stage venture-backed
COGS structure
Dominated by hyperscaler compute/storage for Zilliz Cloud (AWS, GCP, Azure infrastructure underneath the managed service) plus support engineering. Milvus 2.6 cost work (index compressed to 1/32 size via RaBitQ 1-bit quantization, tiered storage, Woodpecker object-storage WAL replacing Kafka/Pulsar dependencies) is explicitly aimed at cutting per-vector serving cost - both a customer price lever and a gross-margin lever.
Capex
Minimal owned capex; infrastructure is rented from hyperscalers and passed through as COGS. Primary investment is R&D headcount across US and China engineering.
Latest earnings
n/a
none published
- Total funding
- ~$113M (last: $60M Series B extension, Aug 2022, Prosperity7 Ventures lead)
- Valuation
- not disclosed at any round
- GitHub stars (Milvus)
- 40,000+ (Dec 2025)
- Enterprise deployments (claimed)
- 10,000+ teams running Milvus in live AI systems
- Founded / HQ
- 2017, Shanghai origin; relocated HQ to San Francisco Bay Area 2022
- Founder-CEO
- Charles Xie (ex-Oracle engineer)
Growth drivers
- RAG and agentic-AI memory workloads driving enterprise vector-search demand
- OSS-to-cloud conversion of the 10,000+ claimed Milvus enterprise deployments (NVIDIA, Salesforce, eBay, Airbnb, DoorDash cited in the Dec 2025 release; Walmart, AT&T, Zillow, Intuit, OpenEvidence, Exa in Zilliz customer rosters)
- Vector Lakebase (public preview Jun 2026) — expansion from point-solution vector DB into a unified lake-native data platform for AI (plus the Loon lake-native storage engine), widening TAM
- Cost/performance leadership claims vs Elasticsearch (3-4x full-text/BM25 throughput, up to 7x on some datasets) and billion-scale serving at lower cost (Milvus 2.6, Jun 2025)
- Enterprise-grade features unlocking regulated buyers — native cross-region disaster recovery (Mar 2026; claimed only vector DB with automated cross-region failover, sub-60s recovery), SOC 2 Type II, 99.95% uptime SLA
- Multi-cloud (30+ regions) + AWS Marketplace distribution reducing procurement friction
Bull & bear
Zilliz owns the deepest-adopted open-source asset in the fastest-growing database category of the AI era, and is executing the proven open-core playbook (MongoDB, Confluent, Databricks) just as enterprise RAG and agent-memory workloads move from pilots to production. If even a modest share of its 10,000+ claimed OSS deployments converts to Zilliz Cloud, and Vector Lakebase lands as a platform, the company compounds into a strategic AI-infrastructure asset with multiple exit paths.
- Category-leading OSS gravity: 40k+ GitHub stars with one of the fastest star-growth spurts in project history in late 2025 - developer mindshare is the top of the COSS funnel and Milvus leads it
- Production-scale differentiation is real: billion-vector serving, 32x index compression (RaBitQ), GPU acceleration, and automated cross-region DR are hard engineering that pgvector-class substitutes cannot match at scale
- Vector Lakebase (Jun 2026) expands TAM from the standalone vector-DB niche toward the far larger AI data-platform market, reframing Zilliz as a data platform rather than an index
- AI-native customer roster (OpenEvidence, Exa, DoorDash, Zillow) shows the product wins where retrieval quality and cost at scale actually matter
- Capital-efficient posture since 2022 (no down round, no distressed raise disclosed) into an AI-infra funding market that rewards category leaders - a future up-round or strategic acquisition are both live paths
- Full-text search performance claims vs Elasticsearch (3-4x BM25 throughput at ~1/3 index size) open a second displacement motion inside AI-native stacks
Vector search is becoming a checkbox feature of every incumbent database and every hyperscaler platform, collapsing the standalone category's pricing power. Zilliz monetizes only when OSS users choose not to self-host and not to use the database they already pay for - a narrowing slice - while its opaque financials, four-year gap since the last disclosed round, and China-linked history add financing and procurement risk that public-market or acquirer diligence will price hard.
- Feature commoditization: pgvector, MongoDB Atlas, Elastic, Redis, and all three hyperscalers now ship good-enough vector search inside databases enterprises already run and already pay for - the standalone vector DB buy decision is disappearing for mid-scale workloads
- Open-core leakage is structural: Milvus's own success is the bear case - the largest, most sophisticated users (the best potential customers) are exactly those most capable of self-hosting free forever
- No disclosed primary round since Aug 2022 and no disclosed valuation ever: if growth were unambiguously strong, an up-round announcement would be cheap signaling; its absence is at minimum uninformative and at worst adverse
- Competitive squeeze from both ends: Pinecone owns zero-ops serverless mindshare; Qdrant/Weaviate/Chroma undercut on price and developer simplicity; Databricks and Snowflake absorb the lakehouse story Vector Lakebase is chasing
- Geopolitical overhang: Shanghai origins, Chinese growth investors, and a Saudi state-linked lead investor create real procurement friction in US government, defense, and some regulated verticals - and would attract CFIUS-style diligence in any strategic exit
- Architecture risk: very-long-context models, model-native retrieval, and agent frameworks with built-in memory could compress the independent retrieval layer before Zilliz reaches escape-velocity scale
- Revenue base (even if the ~$49M 2024 estimate is directionally right, and the other circulating estimate says 2023 was ~$16M) is small relative to the competitive capital arrayed against it - Pinecone alone raised $100M at a $750M valuation in Apr 2023 (~$138M total), and hyperscalers spend more on database R&D in a quarter than Zilliz has raised in its lifetime
What it is worth
No disclosed valuation at any round and no public financials, so any figure is a scenario, not a mark. Anchors: (a) last primary capital $60M Series B extension Aug 2022, valuation undisclosed; (b) unverified third-party revenue estimates (~$16.2M 2023 per GetLatka; ~$48.8M 2024 per UpMarket/Sacra); (c) private comps - Pinecone $750M post-money (Apr 2023, $100M a16z-led round), public COSS/database comps (MDB, ESTC) at roughly 5-8x forward revenue in 2026, AI-infra private premiums running 10-25x ARR for category leaders.
~$250M-$400M
commoditization compresses to 5-8x public-database multiples on estimated ARR, with an illiquidity and geopolitical-diligence discount
~$500M-$900M
10-15x estimated ARR, consistent with Pinecone's 2023 mark and a still-growing but contested category
~$1.0-1.5B+
AI-infra premium (15-25x+ ARR) for the OSS category leader, Vector Lakebase lands, or a strategic acquirer pays for the Milvus ecosystem
If the ~$49M 2024 revenue estimate is directionally right, infrastructure-software multiples imply a wide band; if the ~$16M 2023 estimate is closer to reality, all three scenarios shift down materially. Treat all scenarios as analyst constructs on unverified inputs. Vintage 2026-07. Not financial advice.
SWOT
Strengths
- Most-adopted open-source vector database by GitHub traction and claimed enterprise deployment count - default choice for self-hosted, Kubernetes-native vector search at scale
- Deep distributed-systems engineering moat — billion-scale ANN serving, GPU acceleration (NVIDIA cuVS/CAGRA lineage), quantization (RaBitQ), tiered storage, object-storage WAL (Woodpecker)
- Neutral open-source governance (LF AI & Data Foundation graduated project) reduces vendor lock-in fear and drives bottom-up adoption
- Marquee logos across OSS and cloud — NVIDIA, Salesforce, eBay, Airbnb, DoorDash, Walmart, AT&T, Zillow, OpenEvidence, Exa
- Multi-cloud managed service with differentiated enterprise features (native cross-region DR with automated failover, launched Mar 2026)
Weaknesses
- No new primary capital disclosed since Aug 2022 — either efficient growth or constrained war chest vs Pinecone/Databricks-scale rivals; not verifiable
- Classic open-core leakage — the best engineering teams can run Milvus free forever; conversion to paid cloud is the whole business model
- China-origin history and China-linked cap table (Hillhouse, TrustBridge, 5Y, Yunqi) plus Saudi state-linked lead (Prosperity7) can complicate US government/defense and some enterprise procurement
- Self-hosted Milvus is operationally heavy (DevOps burden) - pushes ease-of-use buyers to Pinecone or pgvector
- No public financials — revenue quality, NRR, and burn are opaque even by private-company standards - circulating third-party estimates disagree widely
Opportunities
- Vector Lakebase — expand from vector search feature to system-of-record data platform for AI, competing for lakehouse budgets
- Agentic AI memory layers - persistent, per-agent vector memory is a new workload class beyond RAG
- OSS install base monetization: 10,000+ claimed enterprise deployments is a large, warm conversion funnel
- Full-text + hybrid search displacement of Elasticsearch in AI-native stacks (Milvus 2.5 hybrid search, 2.6 BM25 benchmarks)
- Acquisition target value for a hyperscaler, Databricks, Snowflake, Confluent, or NVIDIA seeking an AI-data anchor
Threats
- Vector search commoditizing into a feature of incumbent databases: pgvector in Postgres, MongoDB Atlas Vector Search, Elastic, Redis, Oracle 23ai, every hyperscaler
- Pinecone's zero-ops serverless positioning wins the ease-of-use segment; Qdrant/Weaviate undercut on price and simplicity
- Hyperscalers bundling vector search into existing commitments (Amazon OpenSearch, Azure AI Search, Google Vertex/AlloyDB) at marginal cost
- Long-context LLMs and integrated retrieval inside model providers could shrink the standalone vector-DB layer
- US-China tech tension — export controls, procurement bans, or investor-origin scrutiny could impair enterprise and government sales
Moats, dependencies & bottlenecks
Moats
LF AI & Data governance) Largest OSS vector-DB community by stars/deployments; but OSS moats protect adoption, not monetization, and can be forked or commoditized
moderate-strong Quantization (RaBitQ), GPU indexes, tiered storage, cross-region DR are multi-year engineering leads over feature-checkbox rivals
Migrating indexes, schemas, and retrieval tuning is painful, but abstraction layers (LangChain, LlamaIndex) deliberately erode this
Brand/default status among AI-native builders for self-hosted vector search Developer mindshare is real but contested quarterly by Pinecone, Qdrant, Chroma marketing
Dependencies
infrastructure + distribution + competitor Zilliz Cloud runs on and is sold through clouds whose first-party services (OpenSearch, Azure AI Search, Vertex) compete directly
technology partner GPU-accelerated indexing (cuVS/CAGRA collaboration) and NVIDIA as a cited Milvus user; concentration in one accelerator ecosystem
The category's growth assumes external retrieval stays architecturally necessary as context windows grow
product pipeline / governance Community goodwill constrains aggressive license changes (the MongoDB/Elastic SSPL path is costlier for a foundation-governed project)
No disclosed raise since Aug 2022; continued growth investment likely requires new primary capital on undisclosed terms
Anthropic, Cohere, open-source models) complementary input Vector DB value scales with embedding quality; provider-agnostic, so low concentration risk
Advantages
- Most-deployed open-source vector database with 40k+ GitHub stars (Dec 2025) and claimed 10,000+ enterprise teams in production
- Proven at billion-vector production scale where most rivals have not been battle-tested
- Cost-performance engineering lead — index compressed to 1/32 size (RaBitQ), 3-4x Elasticsearch full-text throughput claims, tiered storage economics
- Only vector DB claiming automated cross-region disaster recovery with sub-60s failover (launched Mar 2026) - an enterprise wedge
- Multi-cloud neutrality (AWS, Azure, GCP, 30+ regions) vs hyperscaler-locked alternatives
Weaknesses
- Monetization opacity — no disclosed revenue, NRR, or unit economics; the only revenue figures public are unverified, mutually inconsistent third-party estimates
- Stale cap table narrative — last priced event Aug 2022, valuation never disclosed - hard for buyers of the equity story to anchor
- Geopolitical/procurement friction from Chinese origins and China- and Saudi-linked investors
- Operational complexity of self-hosted Milvus cedes the simplicity-seeking segment to competitors
- Sub-scale capital position relative to the strategic importance of the category it leads
Bottlenecks
- OSS-to-paid conversion rate - the single metric the whole business hinges on, and it is undisclosed
- Enterprise sales and compliance machinery (FedRAMP-class certifications, procurement trust) needed to sell into regulated US buyers given the company's China-origin history
- Capital depth vs better-funded rivals and infinitely-funded hyperscalers
- Ease-of-use gap vs serverless-first competitors for the long tail of smaller workloads
Top signals & trends
Top signals
Developer mindshare accelerating three-plus years into the AI cycle, not decaying
TAM expansion move; also a tell that standalone vector search alone is not a big enough business
Either capital-efficient growth or inability/unwillingness to price a round; unverifiable either way
Secondary liquidity exists but no reliable public mark; treat platform valuation estimates as algorithmic and indicative only
Category commoditization pressure on standalone vendors is the dominant industry headwind
Competing on cost-at-scale is the right axis against both hyperscalers and serverless rivals
Trends
Production workloads need scale, DR, and cost engineering - Zilliz's strengths - rather than demo-grade simplicity
New workload class beyond RAG; per-agent vector memory multiplies index count and query volume
The existential category question; standalone vendors must win on scale/cost or be absorbed
Could thin the standalone retrieval layer, though cost per token still favors external retrieval at scale
Validates Vector Lakebase direction but pits Zilliz against far larger platform players
Ongoing scrutiny of China-origin software vendors in US enterprise and government procurement
Ecosystem & competitor graph
Suppliers feed the company; customers pull from it. Line thickness shows the strength of each tie (supply-chain dependency, customer earnings contribution). Hover to isolate a tie.
Primary cloud substrate + marketplace channel for Zilliz Cloud
Cloud substrate for multi-cloud deployments
Cloud substrate for multi-cloud deployments
GPU acceleration partner (cuVS/CAGRA index work); also a cited Milvus user
Cited Milvus enterprise user (Dec 2025 release)
Cited Milvus enterprise user (Dec 2025 release)
Early large-scale Milvus adopter (Dec 2025 release)
Cited Milvus enterprise user (Dec 2025 release)
Cited Milvus enterprise user (Dec 2025 release)
Cited in Zilliz customer rosters
Cited in Zilliz customer rosters
Cited Zilliz Cloud user
Cited in 2022 customer roster
AI-native medical search unicorn cited as Zilliz Cloud user
AI-native web search API company cited as Zilliz Cloud user
Legal AI platform cited as Zilliz customer (2026 milestone release)
Fully managed serverless vector DB; ~$138M raised, $750M valuation (Apr 2023, a16z-led $100M Series B); owns the zero-ops segment; 2025 press reports it explored a sale
Vector search bundled into the document database enterprises already run
Incumbent search platform with dense-vector and hybrid retrieval; Zilliz benchmarks directly against it
Bundled vector search across AWS data services at marginal cost to committed spend
Default retrieval layer for the Azure-OpenAI enterprise stack
ScaNN-lineage vector search integrated in GCP
Rust-based OSS vector DB, aggressive free tier and price positioning; $28M Series A Jan 2024 led by Spark Capital
OSS vector DB with strong Docker-deployment adoption; $67.7M total raised ($50M Series B Apr 2023, Index Ventures)
Developer-first embedded vector store popular in prototyping; weaker at production scale
Mosaic AI Vector Search inside the lakehouse; direct collision course with Vector Lakebase
Cortex vector functions keep AI retrieval inside the warehouse
AI Vector Search in Database 23ai targets its installed base
Free, good-enough vector search inside the world's default OSS database; the biggest silent share-taker at small-to-mid scale
Vector similarity in the ubiquitous cache layer; low-latency niche
Mainland-China cloud vector offerings compete in Milvus's original home market; named for competitive context only, not as investment calls