Epoch AI · Expert-level questions across fields (HLE)
We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.
How to read this leaderboard
What it measures: Humanity's Last Exam was created by the Center for AI Safety and Scale AI. It consists of 2,500 public questions written by experts across multiple fields and was developed in response to saturation in existing benchmarks.
How to read the score: The percentage (%) of correct answers. It evaluates advanced knowledge and reasoning, and scores vary by model and tool-use conditions.
Caveats: Errors have been identified in questions, answer keys and scoring. Interpret small score differences alongside question flaws, evaluation conditions and statistical uncertainty.
Insights from this ranking
- The highest score is 54.8%, achieved by GPT-6 Astra.
- The gap between 1st and 5th place is 10.5%p.
- OpenAI has the most models among the top 10, with 3.
- Results for this test are public for 30 models.
Expert-level questions across fields (HLE) Epoch AI
This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04
Bars show accuracy (%) and start at 0%.
View table (top 30)
| Rank | Model | Developer | Release date | Score |
|---|---|---|---|---|
| 1 | #1OGPT-6 Astra(unknown thinking) | OpenAI | 2026-09-03 | 54.8% |
| 2 | #2Claude Fable 5.1(xhigh) | Anthropic | 2026-09-01 | 46.5% |
| 3 | #3Gemini 3.1 Pro Preview | 2026-02-19 | 46.4% | |
| 4 | #4Gemini 3.8 Flash(unknown) | 2026-09-02 | 44.5% | |
| 5 | #5OGPT-5.4 Pro | OpenAI | 2026-03-05 | 44.3% |
| 6 | #6Muse Spark | Meta | 2026-04-08 | 40.6% |
| 7 | #7Gemini 3 Pro Preview | 2025-11-18 | 37.5% | |
| 8 | #8OGPT-5.4(xhigh) | OpenAI | 2026-03-05 | 36.2% |
| 9 | #9Claude Opus 4.7(unknown) | Anthropic | 2026-04-16 | 36.2% |
| 10 | #10Claude Opus 4.6(max) | Anthropic | 2026-02-05 | 34.4% |
| 11 | #11OGPT-5 Pro | OpenAI | 2025-10-07 | 31.6% |
| 12 | #12OGPT-5.2(unknown thinking) | OpenAI | 2025-12-11 | 27.8% |
| 13 | #13OGPT-5(high) | OpenAI | 2025-08-07 | 25.3% |
| 14 | #14Claude Opus 4.5(unknown thinking) | Anthropic | 2025-11-24 | 25.2% |
| 15 | #15Kimi K2.5 | Moonshot AI | 2026-01-27 | 24.4% |
| 16 | #16OGPT-5.1(unknown thinking) | OpenAI | 2025-11-13 | 23.7% |
| 17 | #17Gemini 2.5 Pro Preview(Jun 2025) | 2025-06-05 | 21.6% | |
| 18 | #18Oo3(high) | OpenAI | 2025-04-16 | 20.3% |
| 19 | #19OGPT-5 mini(unknown thinking) | OpenAI | 2025-08-07 | 19.4% |
| 20 | #20Gemini 2.5 Pro Exp(Mar 2025) | 2025-03-25 | 18.2% | |
| 21 | #21Oo4-mini(high) | OpenAI | 2025-04-16 | 18.1% |
| 22 | #22Claude Sonnet 4.5(unknown thinking) | Anthropic | 2025-09-29 | 13.7% |
| 23 | #23Gemini 2.5 Flash Preview(Apr 2025) | 2025-04-17 | 12.1% | |
| 24 | #24Claude Opus 4.1(unknown thinking) | Anthropic | 2025-08-05 | 11.5% |
| 25 | #25gemini-2.5-flash-preview-05-20 | 2025-05-20 | 11.0% | |
| 26 | #26Claude Opus 4 | Anthropic | 2025-05-22 | 10.7% |
| 27 | #27Gemini 3.1 Flash-Lite | 2026-03-03 | 8.6% | |
| 28 | #28Zglm-4.5 | Z.ai | 2025-08-03 | 8.3% |
| 29 | #29ZGLM-4.5-Air | Z.ai | 2025-07-20 | 8.1% |
| 30 | #30Oo1 Pro | OpenAI | 2025-03-19 | 8.1% |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings