Artificial Analysis has released a benchmark called 'Search Index' designed to measure the quality, cost, and speed of search API providers in AI agent scenarios. The first batch of test subjects includes seven providers: Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, and Brave.
The test uses a standardized agent setup, with all providers using the same model, GPT-5.6 Luna, and only the search provider is changed. The agent runs on Artificial Analysis's open-source framework Stirrup, performing 25 searches and webpage fetches per task.
The benchmark consists of three equally weighted subtests: DeepSearchQA includes 900 research questions requiring multiple searches; the BrowseComp subset tests 200 hard-to-find facts requiring multi-step browsing; and AA-Omniscience covers 600 questions across 6 knowledge domains. Additionally, a no-tool baseline is set, where the model answers independently, serving as a comparison benchmark.