Skip to content
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
All benchmarks

DeepSearchQA

DeepSearchQA asks questions whose answers are lists: every member of a category, every event matching a set of constraints. Its 900 questions each carry a reference list of answer parts, and a response only counts when it finds all of them without padding the list with extras. We hold the model fixed and vary the search engine, request format, and maximum search budget. Where BrowseComp rewards locating one hidden fact, DeepSearchQA rewards finding the complete answer list.

Last benchmark run Aug 11, 2026, 11:09 AM UTC

Dataset
Search providers
Favicon for Perplexity
Favicon for Parallel
Favicon for openai
Favicon for Exa
4 search providers
compared across configurations
Models
Favicon for anthropic
Favicon for openai
Favicon for openai
Favicon for deepseek
4 models
Claude Opus 5, GPT-5.6 Sol, GPT-5.6 Luna, DeepSeek V4 Flash 0731
Test configurations
1 / 5 / 25
search budgets and plugin mode
Winning resultQuality resultsPrice & speedSearch budgetAll configurationsWhy we run itWhat scores tell youHow tasks are scoredMethodology
01

Winning result

The configuration that scored highest, the breakdown behind that score, and the strongest value and speed alternatives.

Highest quality

Favicon for Perplexity
Perplexity
·25-turn
Favicon for anthropic
Claude Opus 5 · high
76.5%answers matched the full reference set

95% confidence range 72.4–80.2%

Winning result in detail

Complete answers
342 of 447
Incomplete answers
105 of 447
Cost / question
$1.69
Latency / question
1.9m
Best value
$0.091 / question
Favicon for Parallel
Parallel
Favicon for deepseek
DeepSeek V4 Flash 0731 · high
25-turn search budget · 72.0% correct
Fastest strong result
1.6m / question
Favicon for Perplexity
Perplexity
Favicon for openai
GPT-5.6 Luna · xhigh
25-turn search budget · 73.0% correct
02

Quality results

For the selected model, each engine is represented by its highest-scoring configuration. An answer counts only when the complete expected list is found without unsupported extras.

Favicon for Perplexity
Perplexity
Claude Opus 5 · high · 25-turn
76.5%
Favicon for Parallel
Parallel
Claude Opus 5 · high · 25-turn
76.4%
Favicon for openai
OpenAI Native
GPT-5.6 Sol · high · 25-turn
75.0%
Favicon for Exa
Exa
Claude Opus 5 · high · 25-turn
70.5%
03

Price and speed

Compare answer quality with average cost and typical time per question. The Pareto line shows the best quality available at each price or latency level.

Price versus quality

Average cost per question on a logarithmic scale. The line shows the best quality available at each price level.

  • Perplexity
  • Parallel
  • OpenAI Native
  • Exa
Price versus quality, exact configuration values
Search providerModelSearch budgetAnswer qualityCost per questionOn the efficiency and quality frontier
PerplexityClaude Opus 5 · high25-turn76.5%$1.69yes
ParallelClaude Opus 5 · high25-turn76.4%$5.03no
OpenAI Native searchGPT-5.6 Sol · high25-turn75.0%$0.57yes
PerplexityGPT-5.6 Sol · high25-turn73.6%$0.57yes
PerplexityGPT-5.6 Luna · xhigh25-turn73.0%$0.10yes
ParallelDeepSeek V4 Flash 0731 · high25-turn72.0%$0.091yes
ParallelGPT-5.6 Sol · high25-turn71.8%$1.48no
ExaClaude Opus 5 · high25-turn70.5%$1.80no
ParallelGPT-5.6 Luna · xhigh25-turn70.0%$0.11no
PerplexityDeepSeek V4 Flash 0731 · high25-turn69.0%$0.063yes
ExaGPT-5.6 Sol · high25-turn69.0%$0.64no
OpenAI Native searchGPT-5.6 Sol · high5-turn67.0%$0.25no
ExaGPT-5.6 Luna · xhigh25-turn67.0%$0.14no
PerplexityClaude Opus 5 · high5-turn64.7%$0.44no
PerplexityGPT-5.6 Luna · xhigh5-turn60.0%$0.032yes
ExaDeepSeek V4 Flash 0731 · high25-turn60.0%$0.096no
PerplexityGPT-5.6 Sol · high5-turn59.8%$0.25no
ParallelDeepSeek V4 Flash 0731 · high5-turn58.0%$0.035no
ExaGPT-5.6 Sol · high5-turn57.9%$0.24no
ParallelGPT-5.6 Sol · high5-turn56.2%$0.48no
PerplexityDeepSeek V4 Flash 0731 · high5-turn56.0%$0.029yes
ExaClaude Opus 5 · high5-turn55.8%$0.41no
ParallelClaude Opus 5 · high5-turn55.3%$0.99no
ExaGPT-5.6 Luna · xhigh5-turn55.0%$0.043no
ExaDeepSeek V4 Flash 0731 · high5-turn53.0%$0.039no
ParallelGPT-5.6 Luna · xhigh5-turn52.0%$0.038no
PerplexityGPT-5.6 Sol · high1-turn45.1%$0.12no
OpenAI Native searchGPT-5.6 Sol · high1-turn45.0%$0.15no
ExaGPT-5.6 Luna · xhigh1-turn42.0%$0.012yes
ExaGPT-5.6 Sol · high1-turn41.7%$0.13no
ParallelGPT-5.6 Sol · high1-turn40.7%$0.15no
PerplexityClaude Opus 5 · high1-turn38.9%$0.093no
PerplexityGPT-5.6 Luna · xhigh1-turn38.0%$0.010yes
ExaClaude Opus 5 · high1-turn33.3%$0.086no
ParallelGPT-5.6 Luna · xhigh1-turn33.0%$0.012no
ParallelClaude Opus 5 · high1-turn32.4%$0.11no

Latency versus quality

Typical run-level average time per question. The line shows the best quality available at each latency level.

  • Perplexity
  • Parallel
  • OpenAI Native
  • Exa
Latency versus quality, exact configuration values
Search providerModelSearch budgetAnswer qualityTypical timeOn the efficiency and quality frontier
PerplexityClaude Opus 5 · high25-turn76.5%1.9myes
ParallelClaude Opus 5 · high25-turn76.4%2.3mno
OpenAI Native searchGPT-5.6 Sol · high25-turn75.0%2.0mno
PerplexityGPT-5.6 Sol · high25-turn73.6%1.7myes
PerplexityGPT-5.6 Luna · xhigh25-turn73.0%1.6myes
ParallelDeepSeek V4 Flash 0731 · high25-turn72.0%1.6mno
ParallelGPT-5.6 Sol · high25-turn71.8%2.0mno
ExaClaude Opus 5 · high25-turn70.5%2.2mno
ParallelGPT-5.6 Luna · xhigh25-turn70.0%1.7mno
PerplexityDeepSeek V4 Flash 0731 · high25-turn69.0%79syes
ExaGPT-5.6 Sol · high25-turn69.0%2.2mno
OpenAI Native searchGPT-5.6 Sol · high5-turn67.0%64syes
ExaGPT-5.6 Luna · xhigh25-turn67.0%1.8mno
PerplexityClaude Opus 5 · high5-turn64.7%48syes
PerplexityGPT-5.6 Luna · xhigh5-turn60.0%64sno
ExaDeepSeek V4 Flash 0731 · high25-turn60.0%1.5mno
PerplexityGPT-5.6 Sol · high5-turn59.8%64sno
ParallelDeepSeek V4 Flash 0731 · high5-turn58.0%1.9mno
ExaGPT-5.6 Sol · high5-turn57.9%82sno
ParallelGPT-5.6 Sol · high5-turn56.2%66sno
PerplexityDeepSeek V4 Flash 0731 · high5-turn56.0%89sno
ExaClaude Opus 5 · high5-turn55.8%50sno
ParallelClaude Opus 5 · high5-turn55.3%55sno
ExaGPT-5.6 Luna · xhigh5-turn55.0%70sno
ExaDeepSeek V4 Flash 0731 · high5-turn53.0%1.7mno
ParallelGPT-5.6 Luna · xhigh5-turn52.0%68sno
PerplexityGPT-5.6 Sol · high1-turn45.1%59sno
OpenAI Native searchGPT-5.6 Sol · high1-turn45.0%66sno
ExaGPT-5.6 Luna · xhigh1-turn42.0%66sno
ExaGPT-5.6 Sol · high1-turn41.7%88sno
ParallelGPT-5.6 Sol · high1-turn40.7%59sno
PerplexityClaude Opus 5 · high1-turn38.9%20syes
PerplexityGPT-5.6 Luna · xhigh1-turn38.0%55sno
ExaClaude Opus 5 · high1-turn33.3%19syes
ParallelGPT-5.6 Luna · xhigh1-turn33.0%72sno
ParallelClaude Opus 5 · high1-turn32.4%22sno
04

Does more search improve complete answers?

Compare complete-answer rates as the maximum search budget increases. Full lists usually need several searches, so this view shows whether additional search closes more answers.

Search provider1-turn5-turn25-turn
Favicon for Perplexity
Perplexity
45.1%GPT-5.6 Sol · high64.7%Claude Opus 5 · high76.5%Claude Opus 5 · high
Favicon for Parallel
Parallel
40.7%GPT-5.6 Sol · high58.0%DeepSeek V4 Flash 0731 · high76.4%Claude Opus 5 · high
Favicon for openai
OpenAI Native
45.0%GPT-5.6 Sol · high67.0%GPT-5.6 Sol · high75.0%GPT-5.6 Sol · high
Favicon for Exa
Exa
42.0%GPT-5.6 Luna · xhigh57.9%GPT-5.6 Sol · high70.5%Claude Opus 5 · high
05

All search configurations

Every verified configuration for all models. Sort by quality, cost, speed, or question count.

#ModelSearch providerSearch budgetReasoning effort
1
Favicon for anthropic
Claude Opus 5
Favicon for Perplexity
Perplexity
25-turnhigh76.5%$1.691.9m447
2
Favicon for anthropic
Claude Opus 5
Favicon for Parallel
Parallel
25-turnhigh76.4%$5.032.3m445
3
Favicon for openai
GPT-5.6 Sol
Favicon for openai
OpenAI Native
25-turnhigh75.0%$0.572.0m100
4
Favicon for openai
GPT-5.6 Sol
Favicon for Perplexity
Perplexity
25-turnhigh73.6%$0.571.7m450
5
Favicon for openai
GPT-5.6 Luna
Favicon for Perplexity
Perplexity
25-turnxhigh73.0%$0.101.6m100
6
Favicon for deepseek
DeepSeek V4 Flash 0731
Favicon for Parallel
Parallel
25-turnhigh72.0%$0.0911.6m100
7
Favicon for openai
GPT-5.6 Sol
Favicon for Parallel
Parallel
25-turnhigh71.8%$1.482.0m450
8
Favicon for anthropic
Claude Opus 5
Favicon for Exa
Exa
25-turnhigh70.5%$1.802.2m444
9
Favicon for openai
GPT-5.6 Luna
Favicon for Parallel
Parallel
25-turnxhigh70.0%$0.111.7m100
10
Favicon for deepseek
DeepSeek V4 Flash 0731
Favicon for Perplexity
Perplexity
25-turnhigh69.0%$0.06379s100

Why we run this benchmark

Finding one fact and finding every fact are different search problems. A single lucky query can land a BrowseComp answer; a DeepSearchQA answer has to be assembled across searches, with the agent tracking what it has found and what is still missing. That coverage pressure exercises a different part of the search stack: result breadth and deduplication matter as much as ranking.

We run it with the model held fixed because that isolates the variables OpenRouter users actually control. Those are which engine handles the searches, whether search runs as a server tool or a plugin, and how many agent turns the loop is allowed. Those knobs are exactly what you can set on a request today.

What the scores can and can't tell you

These scores compare search configurations, not agent products. The model reads search result excerpts only, with no full-page fetching and no code tools, so absolute numbers sit below published agent leaderboards, which allow both. Compare configurations rather than raw levels.

The all-or-nothing grade makes this suite strict: missing one list member scores the same as missing them all, so scores run low and small engine gaps are noisy. Overlapping confidence intervals are treated as unresolved here, not as proof of equality. A configuration's cost is as real as its score, so read quality and efficiency together.

The questions are public, which means memorization can inflate absolute scores over time; differences between search configurations remain meaningful because every configuration shares the same model.

How tasks are scored

Each task is one question with a reference list of expected answer parts. A judge model checks the response against every expected part and also flags excessive answers — parts the response asserts that are not in the reference. A task is correct only when the judge finds every expected part and finds no excessive answers. The grade is binary with no partial credit, and failed or refused tasks score zero.

reward = all_expected_parts_found ∧ no_excessive_answers   // ∈ {0, 1}
judge  = gemini-2.5-flash, strict json_schema completeness verdict
empty or refused answers skip the judge and score 0

A search configuration is one model, search engine, request format, and maximum search budget. Scores are weighted by the number of questions in each verified run; small smoke tests are excluded.

EX

Real run example

Fixed historical runs, independent of the model filter above. Long results and operational metadata are trimmed; selected queries may be edited or omitted.

Task

Intersect three QS university rankings

Which universities in the top 20 of the QS World University Rankings for 2021 were also in the top 10 US University Rankings and the top 10 (world) Arts and Humanities Rankings? Please only use information from the QS 2021 rankings and list the schools in alphabetical order.

Reference

Columbia University, Harvard University, Princeton University, Stanford University, Yale University

Incorrect
Favicon for Exa
Exa
Claude Opus 4.8 · 1 turn

Final answer

Using the QS 2021 lists I could verify, the intersection is Harvard University and Yale University. Confidence: 55%.

What happened

The run found only two of five universities. It never made a distinct query for the QS USA top 10 and missed Columbia, Princeton, and Stanford, so the complete-list grader failed it.

Methodology

Runs execute on OpenRouter's native benchmark harness against production endpoints. Server-tool configurations set a request-level tool-call budget on the web search server tool. Plugin configurations make one pre-inference search of the question with the web search plugin. Engines use the same default configurations that serve production traffic.

Every run persists its exact model, engine, request format, search budget, cost, and available timing telemetry. Missing configurations stay missing in the comparison table, and absent or zero telemetry is not treated as free or instantaneous performance.