Valumigo

Epoch AI · Expert-level questions across fields (HLE)

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: Humanity's Last Exam was created by the Center for AI Safety and Scale AI. It consists of 2,500 public questions written by experts across multiple fields and was developed in response to saturation in existing benchmarks.

How to read the score: The percentage (%) of correct answers. It evaluates advanced knowledge and reasoning, and scores vary by model and tool-use conditions.

Caveats: Errors have been identified in questions, answer keys and scoring. Interpret small score differences alongside question flaws, evaluation conditions and statistical uncertainty.

Insights from this ranking

  • The highest score is 54.8%, achieved by GPT-6 Astra.
  • The gap between 1st and 5th place is 10.5%p.
  • OpenAI has the most models among the top 10, with 3.
  • Results for this test are public for 30 models.

Expert-level questions across fields (HLE) Epoch AI

This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04

OGPT-6 Astra
54.8%
AnthropicClaude Fable 5.1
46.5%
GoogleGemini 3.1 Pro Preview
46.4%
GoogleGemini 3.8 Flash
44.5%
OGPT-5.4 Pro
44.3%
MetaMuse Spark
40.6%
GoogleGemini 3 Pro Preview
37.5%
OGPT-5.4
36.2%
AnthropicClaude Opus 4.7
36.2%
AnthropicClaude Opus 4.6
34.4%
OGPT-5 Pro
31.6%
OGPT-5.2
27.8%
OGPT-5
25.3%
AnthropicClaude Opus 4.5
25.2%
Moonshot AIKimi K2.5
24.4%
OGPT-5.1
23.7%
GoogleGemini 2.5 Pro Preview
21.6%
Oo3
20.3%
OGPT-5 mini
19.4%
GoogleGemini 2.5 Pro Exp
18.2%

Bars show accuracy (%) and start at 0%.

View table (top 30)
RankModelDeveloperRelease dateScore
1#1OGPT-6 Astra(unknown thinking)OpenAI2026-09-0354.8%
2#2AnthropicClaude Fable 5.1(xhigh)Anthropic2026-09-0146.5%
3#3GoogleGemini 3.1 Pro PreviewGoogle2026-02-1946.4%
4#4GoogleGemini 3.8 Flash(unknown)Google2026-09-0244.5%
5#5OGPT-5.4 ProOpenAI2026-03-0544.3%
6#6MetaMuse SparkMeta2026-04-0840.6%
7#7GoogleGemini 3 Pro PreviewGoogle2025-11-1837.5%
8#8OGPT-5.4(xhigh)OpenAI2026-03-0536.2%
9#9AnthropicClaude Opus 4.7(unknown)Anthropic2026-04-1636.2%
10#10AnthropicClaude Opus 4.6(max)Anthropic2026-02-0534.4%
11#11OGPT-5 ProOpenAI2025-10-0731.6%
12#12OGPT-5.2(unknown thinking)OpenAI2025-12-1127.8%
13#13OGPT-5(high)OpenAI2025-08-0725.3%
14#14AnthropicClaude Opus 4.5(unknown thinking)Anthropic2025-11-2425.2%
15#15Moonshot AIKimi K2.5Moonshot AI2026-01-2724.4%
16#16OGPT-5.1(unknown thinking)OpenAI2025-11-1323.7%
17#17GoogleGemini 2.5 Pro Preview(Jun 2025)Google2025-06-0521.6%
18#18Oo3(high)OpenAI2025-04-1620.3%
19#19OGPT-5 mini(unknown thinking)OpenAI2025-08-0719.4%
20#20GoogleGemini 2.5 Pro Exp(Mar 2025)Google2025-03-2518.2%
21#21Oo4-mini(high)OpenAI2025-04-1618.1%
22#22AnthropicClaude Sonnet 4.5(unknown thinking)Anthropic2025-09-2913.7%
23#23GoogleGemini 2.5 Flash Preview(Apr 2025)Google2025-04-1712.1%
24#24AnthropicClaude Opus 4.1(unknown thinking)Anthropic2025-08-0511.5%
25#25Googlegemini-2.5-flash-preview-05-20Google2025-05-2011.0%
26#26AnthropicClaude Opus 4Anthropic2025-05-2210.7%
27#27GoogleGemini 3.1 Flash-LiteGoogle2026-03-038.6%
28#28Zglm-4.5Z.ai2025-08-038.3%
29#29ZGLM-4.5-AirZ.ai2025-07-208.1%
30#30Oo1 ProOpenAI2025-03-198.1%

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings