
Exa is an AI search engine and API returning semantically ranked web content for AI agents and LLM apps.
Public customers disclosed at the Series B include Cursor, the AI coding agent, alongside top private-equity and consulting firms whose names the company has not publicly named. The pitch positions Exa as the retrieval layer underneath agentic AI workflows that need to ground answers on the open web.
The Series B round of 85 million dollars was led by Benchmark with Peter Fenton joining the board, with participation from Lightspeed Venture Partners, NVentures (NVIDIA), and Y Combinator at a 700 million dollar post-money valuation, following a Series A led by Lightspeed in July 2024.
Exa trains a transformer to predict the next link rather than the next word, producing a neural and embedding-based retrieval model that ranks pages by semantic conceptual match instead of keyword overlap.
This link-prediction substrate underpins the company's developer search API and the Websets no-code research surface, returning ranked URLs and dense, query-specific excerpts rather than blue-link result pages designed for human clicks.
Exa is competing in a rapidly forming category of retrieval infrastructure for AI agents, alongside Perplexity (consumer plus API), Tavily, You.com, Brave Search API, and newer entrants like Parallel Web. Most consumer search incumbents (Google, Bing) primarily expose APIs designed for human result pages rather than dense agent-grade retrieval.
The Series B narrative frames Exa as a vertical bet on training proprietary retrieval models — link-prediction transformers and self-hosted GPU inference — rather than wrapping existing search providers, which the team argues is the only durable substrate for AI agents that need to read the open web at runtime.
Unlike Google and Bing search APIs designed for human result pages, Exa's API is purpose-built for AI agents: ranked URLs plus dense, query-specific excerpts intended to minimize tokens consumed by downstream LLM calls.
The company also runs its own retrieval-tuned GPU cluster — disclosed at the Series B as 144 NVIDIA H200 GPUs and 3,456 CPUs with stated plans to grow it 5x — which lets it serve sustained throughput above 100 queries per second at sub-450 millisecond latency without third-party inference dependencies.