DeepSearchQA asks questions whose answers are lists: every member of a category, every event matching a set of constraints. Its 900 questions each carry a reference list of answer parts, and a response only counts when it finds all of them without padding the list with extras. We hold the model fixed and vary the search engine, request format, and maximum search budget. Where BrowseComp rewards locating one hidden fact, DeepSearchQA rewards finding the complete answer list.
Last benchmark run Aug 11, 2026, 11:09 AM UTC

The configuration that scored highest, the breakdown behind that score, and the strongest value and speed alternatives.
Highest quality
95% confidence range 72.4–80.2%
Winning result in detail

For the selected model, each engine is represented by its highest-scoring configuration. An answer counts only when the complete expected list is found without unsupported extras.
Compare answer quality with average cost and typical time per question. The Pareto line shows the best quality available at each price or latency level.
Average cost per question on a logarithmic scale. The line shows the best quality available at each price level.
| Search provider | Model | Search budget | Answer quality | Cost per question | On the efficiency and quality frontier |
|---|---|---|---|---|---|
| Perplexity | Claude Opus 5 · high | 25-turn | 76.5% | $1.69 | yes |
| Parallel | Claude Opus 5 · high | 25-turn | 76.4% | $5.03 | no |
| OpenAI Native search | GPT-5.6 Sol · high | 25-turn | 75.0% | $0.57 | yes |
| Perplexity | GPT-5.6 Sol · high | 25-turn | 73.6% | $0.57 | yes |
| Perplexity | GPT-5.6 Luna · xhigh | 25-turn | 73.0% | $0.10 | yes |
| Parallel | DeepSeek V4 Flash 0731 · high | 25-turn | 72.0% | $0.091 | yes |
| Parallel | GPT-5.6 Sol · high | 25-turn | 71.8% | $1.48 | no |
| Exa | Claude Opus 5 · high | 25-turn | 70.5% | $1.80 | no |
| Parallel | GPT-5.6 Luna · xhigh | 25-turn | 70.0% | $0.11 | no |
| Perplexity | DeepSeek V4 Flash 0731 · high | 25-turn | 69.0% | $0.063 | yes |
| Exa | GPT-5.6 Sol · high | 25-turn | 69.0% | $0.64 | no |
| OpenAI Native search | GPT-5.6 Sol · high | 5-turn | 67.0% | $0.25 | no |
| Exa | GPT-5.6 Luna · xhigh | 25-turn | 67.0% | $0.14 | no |
| Perplexity | Claude Opus 5 · high | 5-turn | 64.7% | $0.44 | no |
| Perplexity | GPT-5.6 Luna · xhigh | 5-turn | 60.0% | $0.032 | yes |
| Exa | DeepSeek V4 Flash 0731 · high | 25-turn | 60.0% | $0.096 | no |
| Perplexity | GPT-5.6 Sol · high | 5-turn | 59.8% | $0.25 | no |
| Parallel | DeepSeek V4 Flash 0731 · high | 5-turn | 58.0% | $0.035 | no |
| Exa | GPT-5.6 Sol · high | 5-turn | 57.9% | $0.24 | no |
| Parallel | GPT-5.6 Sol · high | 5-turn | 56.2% | $0.48 | no |
| Perplexity | DeepSeek V4 Flash 0731 · high | 5-turn | 56.0% | $0.029 | yes |
| Exa | Claude Opus 5 · high | 5-turn | 55.8% | $0.41 | no |
| Parallel | Claude Opus 5 · high | 5-turn | 55.3% | $0.99 | no |
| Exa | GPT-5.6 Luna · xhigh | 5-turn | 55.0% | $0.043 | no |
| Exa | DeepSeek V4 Flash 0731 · high | 5-turn | 53.0% | $0.039 | no |
| Parallel | GPT-5.6 Luna · xhigh | 5-turn | 52.0% | $0.038 | no |
| Perplexity | GPT-5.6 Sol · high | 1-turn | 45.1% | $0.12 | no |
| OpenAI Native search | GPT-5.6 Sol · high | 1-turn | 45.0% | $0.15 | no |
| Exa | GPT-5.6 Luna · xhigh | 1-turn | 42.0% | $0.012 | yes |
| Exa | GPT-5.6 Sol · high | 1-turn | 41.7% | $0.13 | no |
| Parallel | GPT-5.6 Sol · high | 1-turn | 40.7% | $0.15 | no |
| Perplexity | Claude Opus 5 · high | 1-turn | 38.9% | $0.093 | no |
| Perplexity | GPT-5.6 Luna · xhigh | 1-turn | 38.0% | $0.010 | yes |
| Exa | Claude Opus 5 · high | 1-turn | 33.3% | $0.086 | no |
| Parallel | GPT-5.6 Luna · xhigh | 1-turn | 33.0% | $0.012 | no |
| Parallel | Claude Opus 5 · high | 1-turn | 32.4% | $0.11 | no |
Typical run-level average time per question. The line shows the best quality available at each latency level.
| Search provider | Model | Search budget | Answer quality | Typical time | On the efficiency and quality frontier |
|---|---|---|---|---|---|
| Perplexity | Claude Opus 5 · high | 25-turn | 76.5% | 1.9m | yes |
| Parallel | Claude Opus 5 · high | 25-turn | 76.4% | 2.3m | no |
| OpenAI Native search | GPT-5.6 Sol · high | 25-turn | 75.0% | 2.0m | no |
| Perplexity | GPT-5.6 Sol · high | 25-turn | 73.6% | 1.7m | yes |
| Perplexity | GPT-5.6 Luna · xhigh | 25-turn | 73.0% | 1.6m | yes |
| Parallel | DeepSeek V4 Flash 0731 · high | 25-turn | 72.0% | 1.6m | no |
| Parallel | GPT-5.6 Sol · high | 25-turn | 71.8% | 2.0m | no |
| Exa | Claude Opus 5 · high | 25-turn | 70.5% | 2.2m | no |
| Parallel | GPT-5.6 Luna · xhigh | 25-turn | 70.0% | 1.7m | no |
| Perplexity | DeepSeek V4 Flash 0731 · high | 25-turn | 69.0% | 79s | yes |
| Exa | GPT-5.6 Sol · high | 25-turn | 69.0% | 2.2m | no |
| OpenAI Native search | GPT-5.6 Sol · high | 5-turn | 67.0% | 64s | yes |
| Exa | GPT-5.6 Luna · xhigh | 25-turn | 67.0% | 1.8m | no |
| Perplexity | Claude Opus 5 · high | 5-turn | 64.7% | 48s | yes |
| Perplexity | GPT-5.6 Luna · xhigh | 5-turn | 60.0% | 64s | no |
| Exa | DeepSeek V4 Flash 0731 · high | 25-turn | 60.0% | 1.5m | no |
| Perplexity | GPT-5.6 Sol · high | 5-turn | 59.8% | 64s | no |
| Parallel | DeepSeek V4 Flash 0731 · high | 5-turn | 58.0% | 1.9m | no |
| Exa | GPT-5.6 Sol · high | 5-turn | 57.9% | 82s | no |
| Parallel | GPT-5.6 Sol · high | 5-turn | 56.2% | 66s | no |
| Perplexity | DeepSeek V4 Flash 0731 · high | 5-turn | 56.0% | 89s | no |
| Exa | Claude Opus 5 · high | 5-turn | 55.8% | 50s | no |
| Parallel | Claude Opus 5 · high | 5-turn | 55.3% | 55s | no |
| Exa | GPT-5.6 Luna · xhigh | 5-turn | 55.0% | 70s | no |
| Exa | DeepSeek V4 Flash 0731 · high | 5-turn | 53.0% | 1.7m | no |
| Parallel | GPT-5.6 Luna · xhigh | 5-turn | 52.0% | 68s | no |
| Perplexity | GPT-5.6 Sol · high | 1-turn | 45.1% | 59s | no |
| OpenAI Native search | GPT-5.6 Sol · high | 1-turn | 45.0% | 66s | no |
| Exa | GPT-5.6 Luna · xhigh | 1-turn | 42.0% | 66s | no |
| Exa | GPT-5.6 Sol · high | 1-turn | 41.7% | 88s | no |
| Parallel | GPT-5.6 Sol · high | 1-turn | 40.7% | 59s | no |
| Perplexity | Claude Opus 5 · high | 1-turn | 38.9% | 20s | yes |
| Perplexity | GPT-5.6 Luna · xhigh | 1-turn | 38.0% | 55s | no |
| Exa | Claude Opus 5 · high | 1-turn | 33.3% | 19s | yes |
| Parallel | GPT-5.6 Luna · xhigh | 1-turn | 33.0% | 72s | no |
| Parallel | Claude Opus 5 · high | 1-turn | 32.4% | 22s | no |
Compare complete-answer rates as the maximum search budget increases. Full lists usually need several searches, so this view shows whether additional search closes more answers.
| Search provider | 1-turn | 5-turn | 25-turn |
|---|---|---|---|
| Perplexity | 45.1%GPT-5.6 Sol · high | 64.7%Claude Opus 5 · high | 76.5%Claude Opus 5 · high |
| Parallel | 40.7%GPT-5.6 Sol · high | 58.0%DeepSeek V4 Flash 0731 · high | 76.4%Claude Opus 5 · high |
| OpenAI Native | 45.0%GPT-5.6 Sol · high | 67.0%GPT-5.6 Sol · high | 75.0%GPT-5.6 Sol · high |
| Exa | 42.0%GPT-5.6 Luna · xhigh | 57.9%GPT-5.6 Sol · high | 70.5%Claude Opus 5 · high |
Every verified configuration for all models. Sort by quality, cost, speed, or question count.
| # | Model | Search provider | Search budget | Reasoning effort | ||||
|---|---|---|---|---|---|---|---|---|
| 1 | Perplexity | 25-turn | high | 76.5% | $1.69 | 1.9m | 447 | |
| 2 | Parallel | 25-turn | high | 76.4% | $5.03 | 2.3m | 445 | |
| 3 | OpenAI Native | 25-turn | high | 75.0% | $0.57 | 2.0m | 100 | |
| 4 | Perplexity | 25-turn | high | 73.6% | $0.57 | 1.7m | 450 | |
| 5 | Perplexity | 25-turn | xhigh | 73.0% | $0.10 | 1.6m | 100 | |
| 6 | Parallel | 25-turn | high | 72.0% | $0.091 | 1.6m | 100 | |
| 7 | Parallel | 25-turn | high | 71.8% | $1.48 | 2.0m | 450 | |
| 8 | Exa | 25-turn | high | 70.5% | $1.80 | 2.2m | 444 | |
| 9 | Parallel | 25-turn | xhigh | 70.0% | $0.11 | 1.7m | 100 | |
| 10 | Perplexity | 25-turn | high | 69.0% | $0.063 | 79s | 100 |
Finding one fact and finding every fact are different search problems. A single lucky query can land a BrowseComp answer; a DeepSearchQA answer has to be assembled across searches, with the agent tracking what it has found and what is still missing. That coverage pressure exercises a different part of the search stack: result breadth and deduplication matter as much as ranking.
We run it with the model held fixed because that isolates the variables OpenRouter users actually control. Those are which engine handles the searches, whether search runs as a server tool or a plugin, and how many agent turns the loop is allowed. Those knobs are exactly what you can set on a request today.
These scores compare search configurations, not agent products. The model reads search result excerpts only, with no full-page fetching and no code tools, so absolute numbers sit below published agent leaderboards, which allow both. Compare configurations rather than raw levels.
The all-or-nothing grade makes this suite strict: missing one list member scores the same as missing them all, so scores run low and small engine gaps are noisy. Overlapping confidence intervals are treated as unresolved here, not as proof of equality. A configuration's cost is as real as its score, so read quality and efficiency together.
The questions are public, which means memorization can inflate absolute scores over time; differences between search configurations remain meaningful because every configuration shares the same model.
Each task is one question with a reference list of expected answer parts. A judge model checks the response against every expected part and also flags excessive answers — parts the response asserts that are not in the reference. A task is correct only when the judge finds every expected part and finds no excessive answers. The grade is binary with no partial credit, and failed or refused tasks score zero.
reward = all_expected_parts_found ∧ no_excessive_answers // ∈ {0, 1}
judge = gemini-2.5-flash, strict json_schema completeness verdict
empty or refused answers skip the judge and score 0A search configuration is one model, search engine, request format, and maximum search budget. Scores are weighted by the number of questions in each verified run; small smoke tests are excluded.
Fixed historical runs, independent of the model filter above. Long results and operational metadata are trimmed; selected queries may be edited or omitted.
Task
Which universities in the top 20 of the QS World University Rankings for 2021 were also in the top 10 US University Rankings and the top 10 (world) Arts and Humanities Rankings? Please only use information from the QS 2021 rankings and list the schools in alphabetical order.
Reference
Columbia University, Harvard University, Princeton University, Stanford University, Yale University
Final answer
Using the QS 2021 lists I could verify, the intersection is Harvard University and Yale University. Confidence: 55%.
What happened
The run found only two of five universities. It never made a distinct query for the QS USA top 10 and missed Columbia, Princeton, and Stanford, so the complete-list grader failed it.
Runs execute on OpenRouter's native benchmark harness against production endpoints. Server-tool configurations set a request-level tool-call budget on the web search server tool. Plugin configurations make one pre-inference search of the question with the web search plugin. Engines use the same default configurations that serve production traffic.
Every run persists its exact model, engine, request format, search budget, cost, and available timing telemetry. Missing configurations stay missing in the comparison table, and absent or zero telemetry is not treated as free or instantaneous performance.