Valumigo

Epoch AI · Research reports (DeepResearch Bench)

We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: FutureSearch's Deep Research Bench, listed by Epoch, evaluates the ability to find and synthesize information from the web to answer research questions. It should be distinguished from another benchmark with the same name that evaluates the quality of long reports.

How to read the score: It aggregates scores from 0~1 using task-specific precision, recall, F1, answer correctness, probability estimation error and other measures. Displaying the score as a percentage does not mean it is the overall task completion rate.

Caveats: It uses RetroSearch to search archived web content under fixed agent conditions. Some scoring includes AI-assisted judgments, so interpret small differences alongside the evaluation conditions.

Insights from this ranking

  • The highest score is 55.3%, achieved by Claude Opus 4.6.
  • The gap between 1st and 5th place is 2.7%p.
  • Anthropic has the most models among the top 10, with 6.
  • Results for this test are public for 25 models.

Research reports (DeepResearch Bench) Epoch AI

This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04

AnthropicClaude Opus 4.6
55.3%
AnthropicClaude Sonnet 4.6
54.9%
AnthropicClaude Opus 4.5
54.8%
OGPT-5.5
54.0%
AnthropicClaude Sonnet 4.5
52.6%
AnthropicClaude Opus 4.8
50.2%
GoogleGemini 3 Flash Preview
49.8%
OGPT-5
49.6%
Anthropicclaude-opus-4-1-20250805
48.3%
GoogleGemini 3.1 Pro Preview
47.8%
xgrok-4-0709
47.3%
AnthropicClaude Opus 4
46.8%
Anthropicclaude-sonnet-4-20250514_2K
46.6%
GoogleGemini 3 Pro Preview
46.3%
AnthropicClaude Haiku 4.5
45.5%
Oo3
45.2%
AnthropicClaude 3.7 Sonnet
43.6%
GoogleGemini 2.5 Pro Preview
42.8%
OGPT-5.1
42.8%
Googlegemini-2.5-pro
41.5%

Bars show accuracy (%) and start at 0%.

View table (top 25)
RankModelDeveloperRelease dateScore
1AnthropicClaude Opus 4.6(high)Anthropic2026-02-0555.3%
2AnthropicClaude Sonnet 4.6(high)Anthropic2026-02-1754.9%
3AnthropicClaude Opus 4.5(high)Anthropic2025-11-2454.8%
4OGPT-5.5(high)OpenAI2026-04-2354.0%
5AnthropicClaude Sonnet 4.5(2k thinking)Anthropic2025-09-2952.6%
6AnthropicClaude Opus 4.8Anthropic2026-05-2850.2%
7GoogleGemini 3 Flash PreviewGoogle2025-12-1749.8%
8OGPT-5(low)OpenAI2025-08-0749.6%
9Anthropicclaude-opus-4-1-20250805Anthropic2025-08-0548.3%
10GoogleGemini 3.1 Pro PreviewGoogle2026-02-1947.8%
11xgrok-4-0709xAI2025-07-0947.3%
12AnthropicClaude Opus 4Anthropic2025-05-2246.8%
13Anthropicclaude-sonnet-4-20250514_2KAnthropic2025-05-2246.6%
14GoogleGemini 3 Pro Preview(low)Google2025-11-1846.3%
15AnthropicClaude Haiku 4.5(low)Anthropic2025-10-1545.5%
16Oo3(medium)OpenAI2025-04-1645.2%
17AnthropicClaude 3.7 Sonnet(2k thinking)Anthropic2025-02-2443.6%
18GoogleGemini 2.5 Pro Preview(Jun 2025)Google2025-06-0542.8%
19OGPT-5.1(low)OpenAI2025-11-1342.8%
20Googlegemini-2.5-proGoogle2025-06-0541.5%
21OGPT-5.2(low)OpenAI2025-12-1141.1%
22GoogleGemini 3.1 Flash-LiteGoogle2026-03-0337.3%
23OGPT-5.4 mini(low)OpenAI2026-03-1736.3%
24OGPT-5.4(low)OpenAI2026-03-0535.1%
25DeepSeekDeepSeek-R1(May 2025)DeepSeek2025-05-2835.1%

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks