Skip to content
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
All benchmarks

HLE

Humanity's Last Exam collects expert-written questions at the frontier of human knowledge, spanning mathematics, the sciences, and the humanities. We run its 2,158 text-only questions as a search benchmark: the model answers with live web search rather than from stored knowledge alone, and a judge grades whether its final answer matches the reference. We hold the model fixed and vary the search engine, request format, and maximum search budget. Where BrowseComp stresses persistent multi-step browsing, HLE shows what live search adds to questions that are difficult because they require expert knowledge.

Last benchmark run Aug 11, 2026, 12:00 PM UTC

PaperDataset
Search providers
Favicon for Perplexity
Favicon for Parallel
Favicon for Exa
Favicon for openai
4 search providers
compared across configurations
Models
Favicon for anthropic
Favicon for openai
2 models
Claude Opus 5, GPT-5.6 Sol
Test configurations
1 / 5 / 25
search budgets and plugin mode
Winning resultQuality resultsPrice & speedSearch budgetAll configurationsWhy we run itWhat scores tell youHow tasks are scoredMethodology
01

Winning result

The configuration that scored highest, the breakdown behind that score, and the strongest value and speed alternatives.

Highest quality

Favicon for Perplexity
Perplexity
·25-turn
Favicon for anthropic
Claude Opus 5 · high
77.4%answers matched the reference

95% confidence range 67.4–85.0%

Winning result in detail

Correct answers
65 of 84
Incorrect answers
19 of 84
Cost / question
$0.46
Latency / question
83s
Best value
$0.16 / question
Favicon for Perplexity
Perplexity
Favicon for anthropic
Claude Opus 5 · high
1-turn search budget · 75.3% correct
Fastest strong result
48s / question
Favicon for Perplexity
Perplexity
Favicon for anthropic
Claude Opus 5 · high
1-turn search budget · 75.3% correct
02

Quality results

For the selected model, compare each engine at its highest-scoring configuration. The result measures overall answer correctness, not expertise in individual subjects.

Favicon for Perplexity
Perplexity
Claude Opus 5 · high · 25-turn
77.4%
Favicon for Parallel
Parallel
Claude Opus 5 · high · 25-turn
76.2%
Favicon for Exa
Exa
Claude Opus 5 · high · 25-turn
74.2%
Favicon for openai
OpenAI Native
GPT-5.6 Sol · high · 25-turn
71.1%
03

Price and speed

Compare answer quality with average cost and typical time per question. The Pareto line shows the best quality available at each price or latency level.

Price versus quality

Average cost per question on a logarithmic scale. The line shows the best quality available at each price level.

  • Perplexity
  • Parallel
  • Exa
  • OpenAI Native
Price versus quality, exact configuration values
Search providerModelSearch budgetAnswer qualityCost per questionOn the efficiency and quality frontier
PerplexityClaude Opus 5 · high25-turn77.4%$0.46yes
ParallelClaude Opus 5 · high25-turn76.2%$1.33no
PerplexityClaude Opus 5 · high1-turn75.3%$0.16yes
ExaClaude Opus 5 · high25-turn74.2%$0.67no
PerplexityClaude Opus 5 · high5-turn71.7%$0.37no
ExaClaude Opus 5 · high5-turn71.4%$0.40no
OpenAI Native searchGPT-5.6 Sol · high25-turn71.1%$0.40no
PerplexityGPT-5.6 Sol · high25-turn71.0%$0.33no
ParallelGPT-5.6 Sol · high25-turn71.0%$0.71no
PerplexityGPT-5.6 Sol · high1-turn70.0%$0.11yes
ParallelClaude Opus 5 · high5-turn69.6%$0.75no
PerplexityGPT-5.6 Sol · high5-turn69.0%$0.25no
ParallelGPT-5.6 Sol · high5-turn69.0%$0.46no
ExaClaude Opus 5 · high1-turn68.5%$0.17no
ExaGPT-5.6 Sol · high25-turn68.4%$0.38no
OpenAI Native searchGPT-5.6 Sol · high1-turn66.0%$0.16no
OpenAI Native searchGPT-5.6 Sol · high5-turn66.0%$0.28no
ParallelClaude Opus 5 · high1-turn65.6%$0.21no
ExaGPT-5.6 Sol · high5-turn65.0%$0.28no
ParallelGPT-5.6 Sol · high1-turn64.0%$0.16no
ExaGPT-5.6 Sol · high1-turn64.0%$0.12no

Latency versus quality

Typical run-level average time per question. The line shows the best quality available at each latency level.

  • Perplexity
  • Parallel
  • Exa
  • OpenAI Native
Latency versus quality, exact configuration values
Search providerModelSearch budgetAnswer qualityTypical timeOn the efficiency and quality frontier
PerplexityClaude Opus 5 · high25-turn77.4%83syes
ParallelClaude Opus 5 · high25-turn76.2%1.9mno
PerplexityClaude Opus 5 · high1-turn75.3%48syes
ExaClaude Opus 5 · high25-turn74.2%1.7mno
PerplexityClaude Opus 5 · high5-turn71.7%85sno
ExaClaude Opus 5 · high5-turn71.4%77sno
OpenAI Native searchGPT-5.6 Sol · high25-turn71.1%2.0mno
PerplexityGPT-5.6 Sol · high25-turn71.0%84sno
ParallelGPT-5.6 Sol · high25-turn71.0%85sno
PerplexityGPT-5.6 Sol · high1-turn70.0%44syes
ParallelClaude Opus 5 · high5-turn69.6%89sno
PerplexityGPT-5.6 Sol · high5-turn69.0%68sno
ParallelGPT-5.6 Sol · high5-turn69.0%75sno
ExaClaude Opus 5 · high1-turn68.5%48sno
ExaGPT-5.6 Sol · high25-turn68.4%1.7mno
OpenAI Native searchGPT-5.6 Sol · high1-turn66.0%79sno
OpenAI Native searchGPT-5.6 Sol · high5-turn66.0%77sno
ParallelClaude Opus 5 · high1-turn65.6%54sno
ExaGPT-5.6 Sol · high5-turn65.0%72sno
ParallelGPT-5.6 Sol · high1-turn64.0%61sno
ExaGPT-5.6 Sol · high1-turn64.0%50sno
04

Does more search help on expert questions?

Compare correct-answer rates as the maximum search budget increases. HLE emphasizes expert knowledge, so this view shows how much additional live search contributes.

Search provider1-turn5-turn25-turn
Favicon for Perplexity
Perplexity
75.3%Claude Opus 5 · high71.7%Claude Opus 5 · high77.4%Claude Opus 5 · high
Favicon for Parallel
Parallel
65.6%Claude Opus 5 · high69.6%Claude Opus 5 · high76.2%Claude Opus 5 · high
Favicon for Exa
Exa
68.5%Claude Opus 5 · high71.4%Claude Opus 5 · high74.2%Claude Opus 5 · high
Favicon for openai
OpenAI Native
66.0%GPT-5.6 Sol · high66.0%GPT-5.6 Sol · high71.1%GPT-5.6 Sol · high
05

All search configurations

Every verified configuration for all models. Sort by quality, cost, speed, or question count.

#ModelSearch providerSearch budgetReasoning effort
1
Favicon for anthropic
Claude Opus 5
Favicon for Perplexity
Perplexity
25-turnhigh77.4%$0.4683s84
2
Favicon for anthropic
Claude Opus 5
Favicon for Parallel
Parallel
25-turnhigh76.2%$1.331.9m84
3
Favicon for anthropic
Claude Opus 5
Favicon for Perplexity
Perplexity
1-turnhigh75.3%$0.1648s93
4
Favicon for anthropic
Claude Opus 5
Favicon for Exa
Exa
25-turnhigh74.2%$0.671.7m89
5
Favicon for anthropic
Claude Opus 5
Favicon for Perplexity
Perplexity
5-turnhigh71.7%$0.3785s99
6
Favicon for anthropic
Claude Opus 5
Favicon for Exa
Exa
5-turnhigh71.4%$0.4077s98
7
Favicon for openai
GPT-5.6 Sol
Favicon for openai
OpenAI Native
25-turnhigh71.1%$0.402.0m97
8
Favicon for openai
GPT-5.6 Sol
Favicon for Perplexity
Perplexity
25-turnhigh71.0%$0.3384s100
9
Favicon for openai
GPT-5.6 Sol
Favicon for Parallel
Parallel
25-turnhigh71.0%$0.7185s100
10
Favicon for openai
GPT-5.6 Sol
Favicon for Perplexity
Perplexity
1-turnhigh70.0%$0.1144s100

Why we run this benchmark

HLE sits at the opposite end of the search spectrum from BrowseComp. Its questions were written by subject-matter experts to be unambiguous but extremely hard, so a model's baseline score is mostly a function of what it already knows. Adding web search turns that into a different question: how much expert-level knowledge can a search configuration retrieve on demand?

We run it with the model held fixed because that isolates the variables OpenRouter users actually control. Those are which engine handles the searches, whether search runs as a server tool or a plugin, and how many agent turns the loop is allowed. Those knobs are exactly what you can set on a request today.

What the scores can and can't tell you

These scores compare search configurations, not agent products. The model reads search result excerpts only, with no full-page fetching and no code tools, so absolute numbers sit below published agent leaderboards, which allow both. Compare configurations rather than raw levels.

Because HLE leans on expertise rather than browsing depth, differences between search configurations are smaller here than on BrowseComp: many questions are answered (or missed) the same way at every budget. Engine gaps with overlapping confidence ranges are treated as unresolved here, not as proof of equality. A configuration's cost is as real as its score, so read quality and efficiency together.

We use the text-only subset of the public dataset (multi-modal questions are excluded), so scores are not directly comparable to full-HLE leaderboards. The questions are public, which means memorization can inflate absolute scores over time; differences between search configurations remain meaningful because every configuration shares the same model.

How tasks are scored

Each task is one question with a short reference answer. The model answers in a fixed format (explanation, exact answer, stated confidence), and a judge model grades whether the extracted answer is semantically equivalent to the reference — the same answer-equivalence grading BrowseComp uses. The grade is binary with no partial credit, and failed or refused tasks score zero.

reward = judge(extracted_answer ≡ reference_answer)   // ∈ {0, 1}
judge  = gpt-4.1 at temperature 0, strict json_schema verdict
empty or refused answers skip the judge and score 0

A search configuration is one model, search engine, request format, and maximum search budget. Scores are weighted by the number of questions in each verified run; small smoke tests are excluded.

EX

Real run example

Fixed historical runs, independent of the model filter above. Long results and operational metadata are trimmed; selected queries may be edited or omitted.

Task

Trace a hummingbird tail sesamoid

Hummingbirds within Apodiformes uniquely have a bilaterally paired oval bone, a sesamoid embedded in the caudolateral portion of the expanded, cruciate aponeurosis of insertion of m. depressor caudae. How many paired tendons are supported by this sesamoid bone? Answer with a number.

This question is published verbatim on the official HLE site and in Figure 3 of the official paper. The public sources do not publish its reference answer, so both fresh runs are shown as ungraded.

Official HLE dataset example

Reference

Reference answer withheld.

Ungraded · answer 3
Favicon for Perplexity
Perplexity
Claude Opus 4.8 · 1 turn

Final answer

Based on the anatomy of the hummingbird tail depressor system, this unique sesamoid bone embedded in the cross-shaped aponeurosis of the m. depressor caudae supports three paired tendons.

Exact Answer: 3 Confidence: 45%

What happened

The shallow run made one search and answered 3, but its selected independent source only established the anatomical context, not that count. With no official public reference, the result remains ungraded.

Methodology

Runs execute on OpenRouter's native benchmark harness against production endpoints. Server-tool configurations set a request-level tool-call budget on the web search server tool. Plugin configurations make one pre-inference search of the question with the web search plugin. Engines use the same default configurations that serve production traffic.

Every run persists its exact model, engine, request format, search budget, cost, and available timing telemetry. Missing configurations stay missing in the comparison table, and absent or zero telemetry is not treated as free or instantaneous performance.