Epoch AI · Research reports (DeepResearch Bench)
We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.
How to read this leaderboard
What it measures: FutureSearch's Deep Research Bench, listed by Epoch, evaluates the ability to find and synthesize information from the web to answer research questions. It should be distinguished from another benchmark with the same name that evaluates the quality of long reports.
How to read the score: It aggregates scores from 0~1 using task-specific precision, recall, F1, answer correctness, probability estimation error and other measures. Displaying the score as a percentage does not mean it is the overall task completion rate.
Caveats: It uses RetroSearch to search archived web content under fixed agent conditions. Some scoring includes AI-assisted judgments, so interpret small differences alongside the evaluation conditions.
Insights from this ranking
- The highest score is 55.3%, achieved by Claude Opus 4.6.
- The gap between 1st and 5th place is 2.7%p.
- Anthropic has the most models among the top 10, with 6.
- Results for this test are public for 25 models.
Research reports (DeepResearch Bench) Epoch AI
This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04
Bars show accuracy (%) and start at 0%.
View table (top 25)
| Rank | Model | Developer | Release date | Score |
|---|---|---|---|---|
| 1 | Claude Opus 4.6(high) | Anthropic | 2026-02-05 | 55.3% |
| 2 | Claude Sonnet 4.6(high) | Anthropic | 2026-02-17 | 54.9% |
| 3 | Claude Opus 4.5(high) | Anthropic | 2025-11-24 | 54.8% |
| 4 | OGPT-5.5(high) | OpenAI | 2026-04-23 | 54.0% |
| 5 | Claude Sonnet 4.5(2k thinking) | Anthropic | 2025-09-29 | 52.6% |
| 6 | Claude Opus 4.8 | Anthropic | 2026-05-28 | 50.2% |
| 7 | Gemini 3 Flash Preview | 2025-12-17 | 49.8% | |
| 8 | OGPT-5(low) | OpenAI | 2025-08-07 | 49.6% |
| 9 | claude-opus-4-1-20250805 | Anthropic | 2025-08-05 | 48.3% |
| 10 | Gemini 3.1 Pro Preview | 2026-02-19 | 47.8% | |
| 11 | xgrok-4-0709 | xAI | 2025-07-09 | 47.3% |
| 12 | Claude Opus 4 | Anthropic | 2025-05-22 | 46.8% |
| 13 | claude-sonnet-4-20250514_2K | Anthropic | 2025-05-22 | 46.6% |
| 14 | Gemini 3 Pro Preview(low) | 2025-11-18 | 46.3% | |
| 15 | Claude Haiku 4.5(low) | Anthropic | 2025-10-15 | 45.5% |
| 16 | Oo3(medium) | OpenAI | 2025-04-16 | 45.2% |
| 17 | Claude 3.7 Sonnet(2k thinking) | Anthropic | 2025-02-24 | 43.6% |
| 18 | Gemini 2.5 Pro Preview(Jun 2025) | 2025-06-05 | 42.8% | |
| 19 | OGPT-5.1(low) | OpenAI | 2025-11-13 | 42.8% |
| 20 | gemini-2.5-pro | 2025-06-05 | 41.5% | |
| 21 | OGPT-5.2(low) | OpenAI | 2025-12-11 | 41.1% |
| 22 | Gemini 3.1 Flash-Lite | 2026-03-03 | 37.3% | |
| 23 | OGPT-5.4 mini(low) | OpenAI | 2026-03-17 | 36.3% |
| 24 | OGPT-5.4(low) | OpenAI | 2026-03-05 | 35.1% |
| 25 | DeepSeek-R1(May 2025) | DeepSeek | 2025-05-28 | 35.1% |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks