AI for Search — The Market Radar 153 companies across 24 segments, every row carrying its sourced traction and a citation. Each segment is a layer — open one for where the gaps are. Sources and what this is not
2026-08 vintage 153 of 153 sourced
120
gaps identified— show all
The shape of the market Estimated segment size (USD)
Market size Companies Gaps
What this shows: how big each segment is, per the best published estimate we could source. Lab assistant search is the largest at $43B and 53x the smallest sized segment (Crawler control & licensing). Scopes differ, so treat each bar as its own estimate.
Lab assistant search $43B 25
Answer engines $16B 25
Product search $9.2B 25
Agent identity & payments $8.0B 26
Enterprise search $7.8B 26
Search incumbents $7.5B 26
AI browsers $4.5B 24
Market & people intelligence $4.4B 25
Visibility measurement $4.4B 26
Legal search $2.7B 26
Vector DBs & retrieval $2.7B 25
Clinical search $2.4B 26
Answer monetisation $2.1B 26
Embeddings & rerankers $1.9B 25
Crawl & extract $1.6B 26
Patent search $1.4B 26
GEO execution $886M 24
Crawler control & licensing $816.7M 24 Bar colour = how crowded the segment is saturated contested open
Bar height is compressed (square-root) because the sourced range spans roughly 53× — on a straight scale most segments would be invisible. Read the printed figure, not the height. 6 of 24 segments carry no bar: no credible published figure exists, and none was invented. Segment market sizes are third-party analyst estimates, each measuring a DIFFERENT scope on a different definition — point estimates for context, not a comparable series, and they must not be summed. 18 of 24 segments carry one; where no credible published figure exists the segment carries none rather than a substitute.
Who is building what Every company in the corpus, stacked by segment
153 of 153 companies in scope · 24 of 24 segments shown · layers ordered and sized by company count — the corpus carries no market caps, so a box earns attention by its evidence: open one for the sourced traction.
S1 Answer engines contested 6 5 gaps identified — open for the gaps, incumbents and entry risks. DuckDuckGo US · Privacy search with optional AI — Duck.ai routes anonymised chats to third-party models, Search Assist gives AI summaries, and a separate 'noai' endpoint serves classic links only Genspark US · Launched as an agentic AI search product generating 'Sparkpages' per query; has since pivoted the same agent stack into an AI workspace that executes multi-step tasks Kagi US · Subscription, ad-free search with its own Teclis/TinyGem crawlers blended with third-party indexes, plus an opt-in 'Quick Answer' assistant — the user, not an advertiser, is the payer Liner Asia · Consumer AI answer engine with an academic/citation mode, grown out of a web-highlighting extension; distribution partly via the Samsung Internet browser Perplexity US · Consumer AI answer engine with a cited-sources UI, plus the Comet agentic browser and a shopping/agent layer; sells Pro/Max subscriptions and distributes via handset and carrier deals Perplexity (Comet Plus) US · AI answer engine running a publisher revenue-share programme that pays for citation in search results, traffic through its browser, and content used by its assistant to complete tasks + Gap A paid, high-quality, deliberately AI-free search product DuckDuckGo proved the demand exists — its no-AI page traffic tripled in a day on 28 May 2026 — but serves it as a free side endpoint of a privacy brand; nobody sells it as the product to the cohort most willing to pay for search. + Gap An engine that sells the disagreement Every product here returns one confident paragraph with footnotes; none surfaces where its sources conflict, which source is load-bearing for which sentence, or a calibrated confidence — the complaint users voice most and the feature nobody ships. + Gap Pricing tied to the answered task rather than a copied seat Perplexity Pro, Opera Neon at $19.90 and Polar at $20 all cluster at the same monthly point, with no vendor pricing by value delivered and none publishing what an answered query costs it. + Gap A continuity guarantee for buyers standardising a workflow on an engine Product mortality here is high — OpenAI retired Atlas nine months after launch, Arc was wound down, two of these seven engines pivoted out of consumer answers — and no contract term, escrow or rating prices that risk. + Gap Retrieval sold as a supply layer to other engines You.com's pivot revealed that the paying customer is often the rival engine or agent builder rather than the consumer, and it is close to alone in serving that demand at consumer-engine answer quality. S2 Lab assistant search saturated 7 5 gaps identified — open for the gaps, incumbents and entry risks. Anthropic (Claude web search) US · Grounded web search and browsing inside Claude, positioned around work tasks and coding rather than a consumer query box; no ad-funded search surface Google (AI Mode / AI Overviews) US · Generative answers inside Google Search — AI Overviews above the ten blue links, and AI Mode as a separate conversational/agentic tab now doing booking, calling and generated mini-apps Google (Gemini app) US · Standalone assistant app with grounded web search, competing directly with ChatGPT for the query-box habit rather than defending the results page Meta (Meta AI) US · Assistant embedded in WhatsApp, Instagram and Facebook plus a standalone app, and from June 2026 an 'AI Mode' that searches Meta's own public posts rather than the open web Microsoft (Copilot / Bing) US · Consumer Copilot answers plus conversational overviews inside Bing, and Copilot Search inside Microsoft 365 — distribution comes from Windows, Edge and Office rather than a query-box habit OpenAI (ChatGPT search / Atlas) US · Web search inside ChatGPT plus agentic browsing; shipped the Atlas standalone browser in Oct 2025 and folded it back into the ChatGPT app and a Chrome extension in 2026 xAI (Grok / DeepSearch) US · Grok assistant with DeepSearch and a proprietary X-post search tool; distribution comes from being embedded in the X app rather than a search destination + Gap The trigger mechanic a product mechanism that converts an owned surface into a search habit. Everyone here is testing placement — an app icon, a tab, a chat box — and nobody is testing what makes a user bring a question to that surface, which is why Meta's billion-user installed base does not move share. + Gap Third-party verification for ads inside generated answers Money is already flowing at scale, but there is no viewability standard for a sentence, no brand-safety adjacency check when the ad sits inside a model-written paragraph, and no incrementality test — the whole IAS/DoubleVerify function is absent for this inventory. + Gap A neutral, methodologically published measure of answer-engine share Sensor Tower and Similarweb disagree materially over the same window (ChatGPT 46.4% vs ~53%; Claude 10.3% vs ~9%) and neither publishes its sampling method, while advertisers, publishers and investors allocate against those numbers anyway. + Gap The answer as an addressable, versioned, diffable artefact Nobody can cite what an assistant said on a given date, prove it changed, or diff two versions — the same missing primitive a publisher disputing a misattribution, a brand disputing a claim and a regulator running an audit all need. + Gap A payment rail for content used without a link Every working monetisation model requires the answer engine to own the inventory; the case that actually removes publisher traffic — a lab synthesising from a page nobody visits — is settled bilaterally or not at all, with xAI's entire 2025 data-licensing line at $88M against $6.4B of losses as the scale marker. S3 Regional challengers contested 6 5 gaps identified — open for the gaps, incumbents and entry risks. Alibaba (Qwen app / Quark) Other · Consumer AI assistant with search and agentic commerce; Quark was the rebuilt AI search app and Qwen has become the flagship consumer surface, wired into Alibaba's own shopping and delivery inventory Baidu Other · China's incumbent search engine rebuilding its results page around AI-generated answers and agents, while shifting revenue from search ads to AI cloud and AI-powered services ByteDance (Doubao) Other · China's largest consumer AI assistant with search and agent features, distributed through Douyin; began charging for a subscription tier in 2026 Felo Asia · Multilingual AI answer engine and agent aimed at Japanese, Korean and Chinese queries, where English-first engines translate badly Mistral AI EU · European frontier lab whose Le Chat assistant is the largest EU-owned consumer answer surface, plus models sold for sovereign deployments Naver Asia · Korea's incumbent portal retrofitting its search with AI Briefing summaries and an 'AI Tab' conversational search mode that routes into Naver Map, reservations and shopping inventory + Gap A credible local answer engine for the large mid-sized language markets — Indonesian, Vietnamese, Brazilian Portuguese, Arabic, Hindi. Locale-native retrieval beats a translated English engine wherever it has been tried, but it has only ever happened where an incumbent portal or national champion funded it. + Gap Locale-native retrieval sold as a supply layer rather than a destination Global engines translate badly into CJK and other non-Latin query patterns, and no vendor packages regional index and query understanding as a licensable input to them. + Gap A monetisation model for markets where a consumer paywall fails Doubao's subscription launch is the only large natural experiment so far and it went backwards, leaving every regional engine without a tested route between free-and-loss-making and a copied Western price point. + Gap Sovereign or region-resident crawl and index outside Europe The Ecosia–Qwant joint venture exists because European buyers wanted to stop licensing Bing or Google; no equivalent supply exists for Latin America, the Gulf, Southeast Asia or India despite the same procurement pressure. + Gap Neutral, bookable local inventory a non-incumbent engine can call Restaurant, retail and ticketing supply is reachable today only by the company that already owns it — Alibaba, Naver, Google — so a regional challenger has no way to make its answers transact. S4 Index operators open 6 5 gaps identified — open for the gaps, incumbents and entry risks. Brave (Brave Search / Ask Brave) US · Consumer answer engine built on Brave's own crawl rather than a Bing/Google licence, bundled with the Brave browser and resold as a search API to agent builders Common Crawl Foundation US · Nonprofit maintaining the petabyte-scale open web crawl archive that underpins most LLM pretraining corpora European Search Perspective (Staan) EU · 50/50 Ecosia–Qwant joint venture operating Staan, a European-built search index sold to alternative search engines and AI companies so they need not licence Bing or Google Internet Archive (Wayback Machine) US · Non-profit web crawler and public archive whose snapshots are a primary independent corpus of the historical web Mojeek UK · Independent web search engine running its own crawler (MojeekBot) and index rather than syndicating Google or Bing results OpenWebSearch.eu (Open Web Index) EU · EU-funded open, public web index intended as a shared substrate for European search engines and LLM grounding + Gap A contractual index-freshness guarantee No search or retrieval API sells an SLA on how stale a result may be, so any product whose correctness depends on recency — pricing, availability, breaking events — is building on an unstated assumption. + Gap Index access priced for agents rather than humans Pricing is per query because that is what existed, while an agent consumes an index per task and then caches and redistributes derived answers; caching rights, redistribution rights and per-token value are all contractually undefined. + Gap Region-resident sovereign crawl and index outside Europe Staan exists because European buyers wanted to stop licensing Bing or Google, and the same procurement logic is live in the Gulf, India, Southeast Asia and Latin America with no supplier to buy from. + Gap Withdrawal that propagates Content pulled back — by a publisher opt-out, a licence expiry or a deletion request — cannot be verified out of an index, its embeddings or its downstream derivatives, and no operator sells proof that it happened. + Gap Indemnity for the buyer of crawled data Downstream vendors inherit the legal exposure of how the corpus was assembled, and Common Crawl's position shows that exposure is real; no operator prices or absorbs it. S5 AI browsers saturated 3 5 gaps identified — open for the gaps, incumbents and entry risks. Atlassian (The Browser Company — Dia / Arc) US · Dia puts a chat-over-your-tabs assistant in the browser and is being repointed at SaaS work surfaces; Arc, its predecessor, was wound down Opera (Neon) Other · Agentic AI browser sold as a paid tier rather than a free default, bundling frontier models and task automation on top of Opera's own browser engine work Polar US · AI browser for knowledge work built by an ex-Perplexity Comet engineer — schedules recurring agent tasks against open tabs rather than answering one-off queries + Gap An agentic browser built for the buyer who has actually paid — the enterprise Admin controls over which sites an agent may act on, session governance, credential scoping and an audit trail are absent from every consumer-shaped product here, and Atlassian's $610M is the only price signal the category has produced. + Gap Pricing tied to completed tasks rather than a copied monthly seat Model cost scales with the number of agent steps while every product in the segment charges a flat fee near $20, so the heaviest users are the least profitable and no vendor publishes what a completed task costs. + Gap A replayable record of what the agent did inside a logged-in session Nothing today lets a user or an employer reconstruct which pages an agent visited, what it clicked, what it bought or what it disclosed — the prerequisite for using one anywhere consequential. + Gap Agent capability delivered into the browser the user already has Asking for a browser switch is the most expensive request in consumer software, and OpenAI's retreat from a standalone browser to a Chrome extension is the strongest available evidence that the wrapper, not the capability, is what users refuse. + Gap A continuity commitment for anyone standardising a workflow on a browser agent With a nine-month shutdown as the category's most visible datapoint, no vendor offers a term, an escrow or an export path that makes committing a team's daily workflow a defensible decision. S6 Answer monetisation contested 7 5 gaps identified — open for the gaps, incumbents and entry risks. Koah US · Ad network for conversational AI apps — inserts contextual sponsored units into chat and answer surfaces for developers who cannot run their own sales team Kontext US · Ad server for generative-AI apps that generates a context-relevant ad from the live conversation rather than from behavioural tracking People Inc. US · Publisher group (formerly Dotdash Meredith) rebuilding revenue away from search referral via first-party ad targeting (D/Cipher), events, licensing and subscriptions ProRata.ai US · Attribution engine that apportions AI answer revenue to the publishers whose content the answer used, plus Gist Answers on-site AI search for publishers ProRata.ai (Gist Answers / Gist.ai) US · Answer engine built only on licensed content, plus a white-label AI answer box publishers embed on their own sites, with per-answer attribution driving a 50/50 revenue share on generative ads ProRata.ai (Gist) US · Attribution engine that decomposes a generated answer into contributing sources and splits revenue proportionally, plus Gist Answers — an embeddable on-site AI answer box for publishers Taboola (DeeperDive) US · Publisher-embedded generative answer engine that keeps the reader on the publisher's own site, monetised by Taboola's ad stack; the ad engine behind it is now sold to other AI apps + Gap A payment rail for uncited use by a frontier assistant Both working models require the answer engine to own the inventory; nothing pays a publisher when ChatGPT or Gemini synthesises from their page without a link, which is the exact case removing their traffic. + Gap Independent audit of the attribution split The vendor decomposing an answer into contributing sources is also a party to the revenue share it computes, and no third party verifies either the decomposition or that a signed licence was honoured. + Gap An unbundled or self-hostable publisher answer box Today the answer engine arrives welded to the intermediary's ad demand, so a publisher with its own sales force cannot take the product without also taking the monetisation. + Gap Metering at inference rather than at crawl Payment mechanisms in this market attach to the fetch, while the value is created when the content is used to compose an answer — and nobody meters or prices that moment. + Gap The mid-tail publisher too small for a bilateral licence with a lab, too large not to feel the traffic loss, and served by no vendor that can transact on its behalf at that scale. S7 Crawler control & licensing contested 8 5 gaps identified — open for the gaps, incumbents and entry risks. Akamai US · Edge network that detects AI bot traffic and routes it to partner monetisation rails instead of only blocking it — technical alliances with TollBit (tollbooth) and Skyfire (tokenised pay-per-request) Cloudflare US · Edge network offering AI-crawler control at the request path — per-crawler allow/charge/block, HTTP 402 'pay per crawl' metering, and bot verification for LLM crawlers Created by Humans US · Licensing marketplace for authors and rights holders to sell AI training rights to their books and creative work Human Native UK · AI data licensing marketplace pairing rights-cleared content suppliers with model developers under standardised agreements Human Native (Cloudflare) UK · AI data licensing marketplace matching rights holders with AI developers, turning unstructured multimedia into licensable, rights-cleared datasets RSL Collective US · Open licensing standard (Really Simple Licensing) embedding machine-readable AI usage terms and royalty rates — subscription, pay-per-crawl or pay-per-inference — into robots.txt, plus a nonprofit collective rights organisation to negotiate and collect for small publishers ScalePost US · Content licensing and AI-access brokerage building a library of licensed publisher content for AI companies to pay to access TollBit US · Per-fetch metering and paywall for AI bots — publishers set access terms and rates, TollBit authenticates AI traffic and bills retrievals + Gap Inference-side metering — a per-use meter the model vendor attests to and an independent party audits, the way ad impressions are counted by a measurer rather than the seller. RSL 1.0 specifies pay-per-inference; nothing in the market measures it. + Gap Licence-compliance auditing with canary-content methodology and contractual audit rights publishers signing flat-fee deals today have no technical means to detect whether content was used, how much, or in which products. Structurally cannot be sold by any vendor that also sells to the labs. + Gap Pricing the user's agent separately from the training crawl An assistant fetching one page because a named user asked is one-to-one and consented, but renders no ads; enforcement products collapse it into 'bot'. Skyfire's Know-Your-Agent token already carries the principal, so the distinction is technically available and commercially unused. + Gap Retraction propagation and verification — when a licence ends, a person exercises erasure, or a takedown lands, nobody can confirm the content stopped being retrievable through indexes, caches and derived embeddings. + Gap A self-serve tier for the mid-tail publisher too small for a negotiated deal, large enough to depend on referral. Metering vendors will onboard them but pay in crawl money, which is the wrong unit. S8 Agent identity & payments open 5 5 gaps identified — open for the gaps, incumbents and entry risks. Catena Labs US · Financial rails and identity/governance layer for AI agents — spending limits, approved counterparties, audit trails, plus an agent-identity protocol Coinbase (x402) US · HTTP-402-based machine payment protocol letting an agent pay per request for a resource without a human-held account Skyfire US · Payment and identity network for autonomous agents — verified agent identity ('Know Your Agent') plus a payment token settling inline when an agent requests protected content Stripe (Agentic Commerce Protocol) US · Co-author with OpenAI of ACP, the open standard for agent-initiated checkout, plus the Shared Payment Token that lets an agent initiate payment without exposing the buyer's credentials Visa (Trusted Agent Protocol) US · Open protocol letting merchants cryptographically verify a legitimate shopping agent and separate it from malicious bots at checkout + Gap A neutral agent-identity root of trust with published revocation, governed as a consortium rather than shipped as a product — the piece every private scheme presupposes and none can supply for a rival's network. + Gap An agent-side attribution standard a signed provenance header on the order carrying which assistant, which surface and which citation set drove it. Under ACP the merchant learns what was bought, not what caused it, and no single merchant has the leverage to demand it. + Gap Cross-network honouring of a verification issued elsewhere — a translation and trust-bridging layer so a KYA-style credential presented at one edge provider is accepted at another. + Gap Pricing on the principal rather than the user agent Skyfire's token already carries who the agent acts for, so a per-read settlement that distinguishes a consented human-initiated fetch from bulk crawling is buildable today and unsold. + Gap Merchant-side liability and dispute tooling for agent-initiated purchases — what happens on a wrong order, a refund or a chargeback when the buyer relationship belongs to the assistant. S9 Visibility measurement saturated 9 5 gaps identified — open for the gaps, incumbents and entry risks. Ahrefs Asia · Bootstrapped SEO toolset whose Brand Radar tracks brand mentions and citations across AI Overviews/AI Mode, ChatGPT, Copilot, Gemini, Perplexity and Grok alongside YouTube and Reddit AthenaHQ US · GEO platform tracking brand presence in ChatGPT, Gemini and Perplexity, mapping AI responses to the citation sites that produced them and generating optimisation recommendations Bluefish AI US · Enterprise platform to monitor and influence how LLMs represent a brand across ChatGPT, Meta AI, Google AI and retail assistants such as Amazon Rufus Brandlight Asia · Enterprise AI-visibility platform converting AI-answer signals into recommendations for marketing, PR, content and media teams, extending into AI-native ad placement Evertune US · Statistical brand-measurement product that polls 11 AI systems (ChatGPT, Claude, Perplexity, Gemini, AI Mode/Overviews, Copilot, DeepSeek, Meta AI and others) across thousands of prompt runs to produce a brand index and competitor benchmark Peec AI EU · GEO analytics product tracking brand visibility and citations across ChatGPT, Perplexity, Claude, Gemini and Google AI Overviews, sold to mid-market marketing teams and agencies Profound US · Enterprise platform running prompt panels against ChatGPT, Perplexity, Gemini and AI Overviews to measure brand mention share, citation sources and answer sentiment, plus AI-crawler/agent traffic analytics Semrush (Adobe) US · SEO/competitive-intelligence suite whose AI Toolkit tracks brand mentions and citations in LLM answers alongside classic keyword and backlink data Similarweb Asia · Public digital-intelligence vendor whose clickstream panel is being repurposed to measure AI-assistant traffic, referrals and AI-app usage, and licensed into AI products themselves + Gap A panel-measured prompt-volume dataset built on consented client-side panels — the denominator every vendor currently substitutes with an authored prompt list. Whoever produces a credible one owns the primitive the whole category is priced against. + Gap An accreditation regime or public reproducibility benchmark scoring vendors on rerun stability — the MRC analogue the IAB standard implies but explicitly declines to be. Cannot be sold by anyone who also sells visibility. + Gap An incrementality harness matched prompt cohorts, pre-registered intervention, effect size with confidence intervals. Paid media has holdouts and geo-splits as standard; here every reported gain is confounded with ordinary model variance and weekly model updates. + Gap Remediation when an assistant is wrong about a business — a structured, evidence-backed dispute record a model vendor can act on. Detection of misportrayal is well funded across Profound, Brandlight and Bluefish; correction has no vendor and no process at all. + Gap Non-English, non-US measurement The funded cohort's panels are English-language against US-default model behaviour, and retrieval corpora in other languages are thinner, so the underlying behaviour differs rather than merely translating. S10 GEO execution contested 5 5 gaps identified — open for the gaps, incumbents and entry risks. AirOps US · AI-search content engine — measures AI and organic visibility, then runs agents that create, refresh and distribute content plus off-site brand mentions against the gaps it finds Botify US · Enterprise search-visibility platform that instruments how AI answer engines and shopping agents crawl a site, and serves agent-readable product feeds (AgenticCatalog) Daydream US · Agent-plus-expert service for organic search and AI visibility — content strategy, technical work and AI agents run as a managed engagement rather than a self-serve tool Scrunch AI US · Monitors brand appearance in AI search and serves an 'agent experience' layer — a parallel crawler-readable rendering of the customer's site for LLM crawlers Writesonic US · Former AI-writing tool repositioned as an AI-search growth engine — tracks brand recommendations across ChatGPT, Gemini and Perplexity and publishes large-scale answer-stability studies (161,286-prompt cross-engine citation study; 10.7M tracked answer changes) + Gap An incrementality harness for GEO matched prompt cohorts, staggered rollout or synthetic control, pre-registered intervention and a published effect size. Not one execution vendor can currently separate its work from ordinary model variance. + Gap SMB and local presence The funded cohort is priced and staffed for enterprise brands with content teams; a restaurant, clinic or trades business has the identical job and nothing exists between a free checklist and an enterprise contract. Likely a structured-data-first, low-touch product bundled into tools SMBs already buy, not another dashboard. + Gap Merchant-side readability for shopping agents — machine-verifiable attributes, availability and price for an agent issuing the query on a buyer's behalf, which is a different artefact from a merchandised results page and has no established standard or vendor. + Gap Citation-decay monitoring a page that earned a citation can be edited, paywalled, redirected or deleted the next day, the link still resolves, and nothing watches whether it still supports the claim. + Gap Execution outside US English — the playbooks, prompt sets and content engines are built against US-default assistant behaviour, and the retrieval corpora that determine outcomes differ by language rather than translating. S11 Enterprise search contested 5 5 gaps identified — open for the gaps, incumbents and entry risks. Dashworks (HubSpot) US · Workplace search assistant that answers questions across apps, email, docs and meeting notes Glean US · Permission-aware search and assistant over a company's SaaS apps, plus an agent-building layer on the same index Moveworks (ServiceNow) US · Enterprise AI assistant with cross-system employee search over IT, HR and internal knowledge sources Onyx (formerly Danswer) US · Open-source enterprise search and chat assistant connecting 40+ internal sources, self-hostable Sana (Workday) EU · AI search and agents over a company's data sources (Workday, Google Drive, SharePoint, Office 365), bundled with a learning platform + Gap Permission remediation as a product — find and fix the accumulated access mistakes that will embarrass the customer the day search is switched on. Every vendor's answer is 'we respect your existing ACLs', which is exactly the problem, since the existing ACLs are wrong. + Gap Time-aware retrieval with an explicit 'as of date X' control and validity windows as first-class fields Corpora are full of superseded versions and semantic ranking treats v3 and v7 alike; enterprise RAG benchmarking has found hallucination rates around 29% on temporal questions versus 9% on atemporal ones. + Gap Conflict surfacing instead of silent synthesis When two internal sources disagree on a policy, contract term or clinical protocol, the disagreement is the finding — and every product instead emits one confident paragraph. + Gap Retrieval over non-prose artefacts the margin assumption inside a spreadsheet formula chain, the part revision in CAD and PLM, the decision made on a recording. Connector lists in this segment are lists of text apps. + Gap Expertise location — returning the colleague rather than the file, inferred from commits, tickets, review history and meeting participation that already sit inside the indexes these vendors build. S12 Search incumbents saturated 6 5 gaps identified — open for the gaps, incumbents and entry risks. Amazon Web Services (Kendra / Bedrock Knowledge Bases) US · Managed enterprise search (Kendra, 32+ connectors) now superseded by Bedrock Managed Knowledge Bases, AWS's RAG-native retrieval service Coveo Other · Enterprise relevance platform for site, commerce, service and workplace search, now sold as a generative answering layer over the same index Elastic US · Elasticsearch-based search and vector retrieval platform, repositioned as the 'Search AI' substrate enterprises build RAG on Lucidworks US · Solr-based enterprise and commerce search platform (Fusion), retrofitted with generative answering Mindbreeze US · Enterprise AI search appliance and cloud (InSpire) indexing structured and unstructured systems with permission-aware retrieval Sinequa (ChapsVision) EU · Long-standing enterprise search platform for regulated large enterprises, now sold as a RAG/generative-answer layer inside ChapsVision + Gap Validity-window retrieval at the substrate layer — an 'as of' filter and document validity period as first-class index fields, so a superseded version is ranked as superseded rather than as similar text. + Gap Retained, replayable retrieval provenance as a compliance artefact content-addressed snapshots of every retrieved passage, an immutable log and replay of the ranked context that produced an answer. Grounding vendors return a citation to a live URL, which is mutable and therefore not evidence nine months later. + Gap Entitlement and metering plumbing for licensed corpora inside RAG — enforcing per-seat rights, metering publisher usage and attributing revenue once retrieved passages flow into a generated answer. Today this is negotiated bilaterally, one contract at a time. + Gap Continuous retrieval-quality drift monitoring connectors break, permission changes silently drop a corpus, an index falls behind, a model swap changes ranking, and customers find out from user complaints. Observability vendors watch latency and cost, not answer quality. + Gap The warehouse-plus-document join The most common real enterprise question needs one fact from a database and one from a deck or thread; text-to-SQL and RAG are sold separately, to different buyers, with the join left to the customer. S13 Legal search saturated 6 5 gaps identified — open for the gaps, incumbents and entry risks. Clio (vLex) Other · Legal research platform (vLex) with the Vincent AI assistant, now folded into Clio's practice-management system Harvey US · Legal research, drafting and document-review agents over firm and public legal corpora Legora EU · Collaborative legal AI workspace: research, review and drafting across a firm's own documents and legal sources LexisNexis (Lexis+ with Protégé) US · The other legal-research incumbent: Shepard's-backed case-law search with Protégé as the generative research and drafting assistant Luminance UK · Legal-domain LLM that reads, searches and negotiates contracts across a corporate document estate Thomson Reuters (Westlaw / CoCounsel) Other · The legal-research incumbent: Westlaw's licensed case-law corpus with CoCounsel (acquired Casetext, 2023) as the citation-grounded generative layer on top + Gap A vendor-neutral citation verifier that takes a finished brief or memo and resolves every cite against the authoritative corpus with a pass/fail — the research vendors have no incentive to ship a tool that audits their own output, and the sanctionable party is the one holding the risk + Gap Grounded legal research priced for the solo practitioner, the two-partner boutique and the community clinic; seat pricing set for the Am Law 100 leaves the majority of practitioners on general-purpose chatbots with no corpus + Gap Jurisdiction-specific corpora outside US/English law — each is a separate licensing and citator problem no US-scale player finds worth solving, and Legora, the only non-US-headquartered player here, is spending its Series D on US expansion + Gap As-of retrieval which version of the statute, regulation or policy was in force on the date of the conduct. Ranking is semantic, so superseded and current text are equally retrievable and get spliced into one answer + Gap Conflict surfacing rather than silent resolution — when two authorities disagree, the disagreement is the finding, and every product in this segment instead synthesises one confident paragraph S14 Patent search contested 6 5 gaps identified — open for the gaps, incumbents and entry risks. Clarivate (Derwent) UK · The patent-search incumbent: Derwent World Patents Index plus Derwent AI Search and Patent Monitor over curated, human-abstracted patent records DeepIP US · AI copilot for patent attorneys inside Microsoft Word covering drafting, patentability and prior-art search, and office-action responses IPRally EU · Graph-based, explainable patent search — parses claims into technical knowledge graphs rather than embedding whole documents Patlytics US · AI platform for the patent lifecycle — prior-art and invalidity search, claim charts, drafting and portfolio analysis PatSnap Asia · IP and innovation-intelligence platform with agentic patent search, drafting and landscape analytics across patents, litigation and scientific literature Solve Intelligence US · AI platform for patent attorneys covering invention capture, drafting, prosecution and Charts for infringement/invalidity and freedom-to-operate analysis + Gap A public recall benchmark scored against examiner- and court-cited references, run by a party that does not sell search — whoever builds it decides whether AI-native retrieval actually beats Boolean, which is why no vendor will + Gap CJK-first prior art a large share of art originates in Chinese, Japanese and Korean filings and is retrieved today through machine-translated abstracts, with no vendor selling native-language claim parsing as the primary path rather than a fallback + Gap Non-patent literature as first-class art — standards documents, datasheets, conference proceedings, product manuals and public code repositories are what actually invalidates claims and sit outside the curated patent index entirely + Gap Freedom-to-operate as continuous monitoring against a product specification, rather than a one-off search that goes stale the week after it is delivered + Gap Prior-art search priced and packaged for the small IP boutique and the single in-house counsel, who cannot buy a full lifecycle suite but carry the same novelty question S15 Science search contested 5 5 gaps identified — open for the gaps, incumbents and entry risks. Allen Institute for AI (Asta) US · Open scientific-research agent ecosystem — literature discovery and synthesis over the Semantic Scholar corpus, plus AstaBench for measuring science agents Consensus US · AI search engine over scientific papers that answers questions with a claim-level evidence meter tied to real studies Elicit US · AI research assistant that finds papers semantically and automates screening and data extraction for systematic literature reviews scite US · Citation-context search — classifies whether a citing paper supports, contradicts or merely mentions the cited claim Undermind US · Agentic literature search that iteratively expands a citation graph to maximise recall of niche papers, rather than one-shot semantic retrieval + Gap Retrieval into methods and limitations sections rather than abstracts and conclusions — the disconfirming detail sits behind the paywall, and every abstract-level answer is systematically more confident than the paper it cites + Gap A retraction- and replication-aware answer layer that knows at answer time whether a cited paper has been retracted, corrected, or failed a replication attempt, instead of citing it cleanly + Gap Literature search for fields whose real corpus is preprints, code and datasets rather than journals — citation graphs are structurally stale relative to arXiv-speed fields, so graph-expansion recall degrades exactly where the research moves fastest + Gap An audited claim-verification API sold per verified claim to other products, rather than another destination search site competing with a free bundled assistant mode + Gap Grey and non-English literature — theses, agency reports, trial and study registries, regional journals — which never enters the indexed citation graph and therefore cannot be contradicted by it S16 Clinical search contested 5 5 gaps identified — open for the gaps, incumbents and entry risks. Atropos Health US · Query layer over de-identified real-world patient data — ChatRWD turns a clinical question into a publication-grade observational study Elsevier (ClinicalKey AI) EU · Clinical decision-support search answering point-of-care questions over full-text journals and practice guidelines with source citations OpenEvidence US · Point-of-care clinical search and answer engine for physicians, grounded in licensed medical-journal content rather than the open web; monetised by pharma advertising, free to clinicians Pathway (Doximity) Other · AI clinical reference that answers point-of-care medical questions from a structured corpus of guidelines, drug labels and landmark trials Wolters Kluwer (UpToDate) EU · The clinical-reference incumbent: editorially-authored, graded evidence summaries sold per clinician seat, with UpToDate Expert AI as the generative layer + Gap Independent, standing measurement of advertiser influence in ad-funded clinical answer engines — a public methodology run by a party that takes no pharma money, not a vendor's own disclosure page + Gap National and regional guideline search outside the US guidelines are set nationally and licensed per country, so each market is a separate corpus problem no US-scale player finds worth solving + Gap Evidence search for the non-physician clinical workforce — nurses, pharmacists, allied health, care managers — whose questions, formulary constraints and liability differ from a physician's and who are not the target of a physician-verified product + Gap Patient- and caregiver-facing evidence search over the same grounded corpus with reading-level control; the current alternative is a general chatbot with no licensed clinical content behind it + Gap Answers that surface disagreement between guideline bodies and flag what changed since the last revision, rather than synthesising one confident paragraph out of sources that conflict S17 Market & people intelligence contested 5 5 gaps identified — open for the gaps, incumbents and entry risks. AlphaSense US · Market-intelligence search over filings, transcripts, broker research and Tegus expert-call transcripts, with generative summarisation on top Clay US · Go-to-market data platform whose Claygent research agents search the web and 150+ data providers to enrich company and people records Hebbia US · Matrix — a grid-based retrieval and extraction workspace that runs one question across thousands of documents (filings, diligence rooms) and returns cited cells Juicebox US · PeopleGPT — natural-language search over ~800M professional profiles stitched from LinkedIn, GitHub, Google Scholar and personal sites, plus outreach agents Rogo US · Agentic research system for finance that searches filings, transcripts and a firm's own data warehouse from inside Excel, PowerPoint and Word + Gap A consent-based or properly licensed data supply for people search that a risk-averse enterprise buyer can purchase instead of scraped profiles — the entire segment currently runs on provenance that ranges from public-web scraping to ambiguous + Gap Cross-organisation federated search with contractual boundaries — diligence, joint ventures, outside counsel, supply chains — where each side's visibility and retention limits are enforced by the system instead of by copying documents into a data room and losing provenance + Gap The join between a warehouse number and a document reason ('why did EMEA churn spike in Q2') text-to-SQL and document retrieval are sold by different vendors to different buyers, and the most common real question crosses that line + Gap Private-company and non-English coverage — the licensed corpora that make market intelligence valuable are filings and English-language broker research, leaving mid-market private companies and non-US-language sources thin + Gap Expertise location inside the enterprise — who has done this deal type, debugged this system — inferable from signals already in the index but sold by nobody as a retrieval product S18 Product search saturated 5 5 gaps identified — open for the gaps, incumbents and entry risks. Algolia US · Hosted search-and-discovery API for websites and apps, with vector/hybrid retrieval and merchandising controls Amazon (Rufus) US · Conversational product-discovery assistant embedded in Amazon's storefront, answering shopping questions over catalogue, reviews and Q&A Bloomreach US · Commerce search and product-discovery platform; Loomi Connect exposes its product discovery over MCP so assistants like ChatGPT can query a retailer's catalogue Constructor US · Clickstream-trained e-commerce search and product discovery optimised against revenue-per-visitor rather than relevance scores Vantage Discovery (Shopify) US · Generative product-discovery search over merchant catalogues, built by ex-Pinterest search engineers + Gap The merchant side of agentic commerce making a catalogue machine-verifiable — attributes, real-time availability, price, returns terms — for buying agents, with no established standard and no vendor selling it as a product + Gap Attribution for agent-mediated purchases a merchant cannot tell which agent, prompt or answer produced a sale, so no budget can be allocated against the channel that is taking the intent + Gap Discovery for considered, spec-driven buying — industrial parts, components, compatibility and fitment — where the job is constraint satisfaction against a specification, not ranked relevance over a text index + Gap The long tail of small and local merchants, whose integration budget and catalogue size are nothing like the enterprise deployments these APIs are priced and engineered for + Gap Ranking optimised on margin after returns rather than revenue per visitor — the metric every vendor sells against is the one that overstates the value of a high-return conversion S19 Agent search APIs contested 8 5 gaps identified — open for the gaps, incumbents and entry risks. Brave (Brave Search API) US · Independent web index sold as a developer API with Search, Answers and an LLM-Context endpoint tuned for RAG and agent grounding Exa US · Neural web search index and API returning full page contents and structured result sets for LLMs, plus Websets for high-compute enumerative search Linkup EU · Web search API for AI apps with Search, Fetch and Research endpoints, letting developers tune the index, restrict sources and set content freshness Microsoft (Grounding with Bing Search) US · Azure AI Foundry Agent Service tool that hands web-derived context to an agent, replacing the retired standalone Bing Search API Parallel Web Systems US · Web research and search API for AI agents — Task and Search APIs plus Monitors that watch pages for meaningful change and chain follow-up tasks Perplexity (Sonar / Search API) US · Sells its answer-engine stack to developers — a Search API returning ranked web results, an Agent API with built-in web-search and URL-fetch tools, and Sonar search-grounded chat models Tavily Asia · Search API purpose-built for agents that returns pre-scraped, pre-ranked text chunks with citations instead of link lists You.com US · Search and research APIs (web search, news, research agents) sold as the retrieval layer inside other companies' AI products, alongside its own assistant + Gap No vendor sells recall as a contractual number on enumerative search — Websets and Parallel Tasks return a set with no stated coverage bound, so a buyer running competitive or compliance enumeration cannot tell what was missed + Gap Freshness is a tuning knob (Linkup) but never a guarantee; nobody prices a per-source fetch-age SLA for agents acting on time-sensitive facts + Gap Continuous change-detection with a durable replayable audit trail exists only as Parallel Monitors — no one serves regulated monitoring (sanctions, label changes, disclosure updates) where the record of what the web said when is the product + Gap No search API returns the licensing status of the passage it hands back, leaving the buyer to guess whether a returned chunk can legally reach a commercial output + Gap Auditable non-English and per-country index depth all four are anglophone-web-first, and no vendor publishes coverage by language or jurisdiction that a buyer could verify S20 Crawl & extract contested 11 5 gaps identified — open for the gaps, incumbents and entry risks. Apify Other · Marketplace and cloud runtime for 'Actors' — packaged scrapers and web-automation programs — now exposed to agents over MCP and the x402 payment protocol Bright Data Asia · Proxy network, web-scraping APIs and refreshed web datasets, now sold as web access and data infrastructure for AI labs and agents Crawl4AI Other · Apache-2.0 open-source crawler that renders pages and emits LLM-ready markdown or LLM/CSS-extracted structured data, with a Docker API server and a hosted cloud Decodo (formerly Smartproxy) Other · Proxy pools plus an all-in-one web-scraping API with an AI parser that formats extracted data Diffbot US · Computer-vision-based Extract and Crawl APIs plus a Knowledge Graph of the public web queryable through a DQL search API Firecrawl US · Crawl/scrape/extract API that converts websites into markdown or structured JSON for AI apps, with an open-source core Jina AI EU · Reader (r.jina.ai) converts any URL — including PDFs, Office docs and images — into LLM-ready markdown, with s.jina.ai for search, alongside embedding and reranker models Kadoa EU · LLM-driven extraction platform that turns unstructured web and document sources into schema-conforming data feeds, aimed at financial and retail-intelligence buyers Oxylabs Other · Proxy network, web-scraping API and web-intelligence datasets, repositioned as real-time web-data infrastructure for AI agents Reducto US · Vision-model document parsing and structured-extraction APIs (parse, extract, split, classify) that turn PDFs into retrieval-ready data Unstructured US · Open-source and hosted ingestion pipeline that turns PDFs, images, audio and 20+ other formats into retrieval-ready chunks for RAG + Gap No extraction API returns a rights or licence field with the content, so every buyer re-does the same 'can I use this' analysis by hand downstream + Gap Structured extraction is sold without a measured accuracy or schema-conformance guarantee — buyers hand-audit JSON output with no vendor-published error rate to plan against + Gap Bring-your-own-credential crawling of sources a buyer already licenses (paid databases, subscription trade press, portals) with a compliant audit trail is unserved + Gap Long-tail non-HTML at crawl scale — PDF tables scanned filings, spreadsheets behind portals — is still bespoke work rather than a metered endpoint + Gap Agents have no trustworthy pre-flight cost estimate or hard budget ceiling per crawl, so autonomous crawling is capped by fear of the bill rather than by usefulness S21 Agent browsers contested 7 5 gaps identified — open for the gaps, incumbents and entry risks. Anchor Browser US · Cloud browser for AI agents, with managed authentication and session handling aimed at reliable enterprise web automation Browser Use US · Open-source library plus hosted cloud that lets LLM agents drive a real browser by reading a text representation of the page rather than screenshots Browserbase US · Cloud platform to run, manage and observe headless browser sessions that AI agents drive, plus the Stagehand agent-browser SDK Kernel US · Sandboxed cloud Chromium browsers for agents, provisioned by API in milliseconds, with live view, replay and scoped user-auth grants Lightpanda EU · Headless browser written from scratch in Zig for machines rather than humans — no graphical composition, CDP-compatible, aimed at agent and scraping workloads Skyvern US · Open-source browser-automation agent that uses vision LLMs on rendered pages instead of CSS/XPath selectors, sold as cloud and self-hosted Steel (Nen Labs) Other · Open-source browser API for running fleets of cloud browser sessions that AI agents drive + Gap Per-site reliability is nobody's contract — no vendor publishes or guarantees task success rates on the specific destinations (airlines, banks, government portals) where agent workflows actually fail + Gap Credentialed sessions with custody-grade handling of a user's logins, MFA and consent are unserved, which is precisely what blocks agents from the high-value logged-in web + Gap Pricing is per session-hour rather than per successful task, so the vendor carries none of the failure risk and the buyer cannot forecast cost per completed job + Gap A replayable, tamper-evident action record — what the agent clicked, saw and submitted — is missing for regulated workflows where an auditor, not a developer, is the reader + Gap Nothing signals authorised-agent status to the destination site, so legitimate agents share fate with scrapers and pay the anti-bot tax S22 Code context contested 10 5 gaps identified — open for the gaps, incumbents and entry risks. Anysphere (Cursor) US · AI code editor whose agents run retrieval over an indexed repository to assemble context before editing Augment Code US · Codebase Context Engine that indexes large repos and feeds retrieved context to its coding agents, CLI and IDE plugins Cognition (DeepWiki) US · DeepWiki generates a navigable AI wiki and Q&A layer over any public GitHub repo, with an MCP server that serves that repo context to coding agents Context7 (Upstash) US · MCP server and API that injects up-to-date, version-specific library documentation and code examples into coding-agent prompts Greptile US · Repo-scale code indexing and retrieval engine, productised as an AI code reviewer that reasons over a full codebase graph Mintlify US · Documentation platform that publishes an MCP server and llms.txt from a customer's docs so coding agents retrieve current API reference instead of guessing Qodo (formerly CodiumAI) Asia · Retrieval-augmented code platform that indexes a repo and feeds context to agents for generation, test writing and review across IDE, git and CLI Sourcegraph US · Cross-repo code search and code-intelligence platform for large enterprises, now sold as the retrieval layer coding agents call — Deep Search, an MCP server and the Amp agent Unblocked Other · Context engine that joins a codebase to the discussions and decisions around it (GitHub, Slack, Jira, Confluence) and serves that context to developers and coding agents Vercel (Grep) US · Grep.app, a fast regex/literal code-search engine over public git repositories, now run inside Vercel and exposed to agents via an MCP server + Gap Version-pinned documentation for a company's own internal libraries and services — Context7 solves this for public packages, and nobody solves it behind the firewall where the hallucinated API costs the most + Gap Per-user permission-aware retrieval for agents, so an agent acting for a contractor sees exactly what that contractor may see, is largely unbuilt outside the largest enterprise platform + Gap Cross-repo blast-radius answers ('who calls this across 400 services, and what breaks') priced per answer rather than per seat + Gap There is no accepted benchmark for whether retrieval gave the agent the right context, so buyers cannot compare vendors and default to whatever their agent ships with + Gap Incremental index cost at monorepo scale on a fast-moving trunk is opaque, leaving large buyers unable to model spend before committing S23 Vector DBs & retrieval saturated 7 5 gaps identified — open for the gaps, incumbents and entry risks. LanceDB US · Columnar multimodal lakehouse (Lance format) with built-in vector and full-text search over object storage MongoDB (Atlas Vector Search) US · General-purpose document database with native vector search plus embedded Voyage AI embedding/reranking Pinecone US · Fully managed serverless vector database for similarity search and RAG retrieval Qdrant EU · Open-source vector search engine with composable dense/sparse/multivector retrieval, sold as managed cloud Vectara US · Managed RAG and agent platform running hallucination detection, citation tracking and correction inside the generation pipeline rather than as post-processing; publishes the open-source HHEM evaluation model and a public hallucination leaderboard Vespa.ai Other · Yahoo-spinout serving engine combining vector, lexical and structured search with machine-learned ranking and inference in one system Zilliz (Milvus) US · Company behind the open-source Milvus vector database, sold as the managed Zilliz Cloud + Gap Retrieval quality as a purchasable, continuously measured product — Vectara's HHEM leaderboard is the only widely-cited independent grounding benchmark in the roster, and no vendor sells ongoing quality measurement against a customer's own corpus + Gap Pricing is per vector, per node or per hour; nobody prices per relevant answer, which is the only unit a RAG buyer can tie to value + Gap Multi-tenant, permission-aware retrieval where the access check happens inside ranking rather than as a post-filter that quietly wrecks recall + Gap Verifiable deletion and freshness — proving a document and its derived embeddings are gone within a stated window is a live compliance need with no standard vendor answer + Gap Agent memory is a different workload from RAG (write-heavy, decaying, long-lived, recency-weighted) and is still being served by databases designed for read-mostly corpora S24 Embeddings & rerankers saturated 5 5 gaps identified — open for the gaps, incumbents and entry risks. Cohere Other · Enterprise model provider whose Embed and Rerank models are a default retrieval layer for private-deployment RAG Contextual AI US · RAG platform shipping a standalone instruction-following reranker and a grounded language model for retrieval-heavy enterprise search TwelveLabs US · Video-understanding models with an Embed API producing multimodal embeddings for semantic search across video, audio, image and text Voyage AI US · Domain-specialised embedding and reranking models for retrieval over legal, financial, code and enterprise text ZeroEntropy US · Reranking and embedding models (zerank, zembed) plus context compression and query rewriting sold as a retrieval-accuracy layer + Gap Rerankers tuned on a customer's own corpus with a reproducible, published evaluation against their queries — buyers currently pick by public leaderboard and discover fit in production + Gap Embedding-version migration is an unowned pain when a model improves, re-embedding a large corpus, running both indexes and proving the swap did not regress retrieval is a project every buyer improvises + Gap Reranking is a per-query tax at agent volumes, and no vendor sells a latency and cost-tiered cascade that degrades gracefully under load instead of a single model at a single price + Gap Auditable per-language retrieval quality for low-resource languages, which is exactly where sovereign and public-sector buyers must justify a choice and cannot + Gap Redaction- and permission-preserving embeddings, so sensitive spans do not become retrievable through the vector even when the source document is access-controlled Point-in-time snapshots, not market censuses. Absence from a radar is not evidence a company does not exist, and funding and traction age fast — re-verify a row against its source before it informs a decision. Not investment, commercial or legal advice.