Valumigo

Epoch AI · PhD-level science questions (GPQA Diamond)

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: GPQA Diamond consists of 198 graduate-level multiple-choice questions in biology, chemistry and physics. Domain experts wrote and reviewed the questions, which are designed to be difficult for non-experts even with internet access.

How to read the score: The percentage (%) of correct answers. Choosing uniformly at random from 4 options gives an expected accuracy of 25%.

Caveats: As leading models approach a perfect score, this test alone may become less useful for distinguishing capabilities. It evaluates scientific knowledge and reasoning, which may differ from everyday work capabilities.

Insights from this ranking

  • The highest score is 95.8%, achieved by GPT-6 Astra.
  • The gap between 1st and 5th place is 1.0%p.
  • OpenAI has the most models among the top 10, with 4.
  • Results for this test are public for 30 models.

PhD-level science questions (GPQA Diamond) Epoch AI

This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04

OGPT-6 Astra
95.8%
AnthropicClaude Sonnet 5.5
95.6%
OGPT-6.1 Sol
95.4%
GoogleGemini 3.8 Flash
95.4%
GoogleGemini 3.7 Flash
94.8%
OGPT-5.4 Pro
94.6%
GoogleGemini 3.1 Pro Preview
94.4%
OGPT-6 Sol
94.3%
GoogleGemini 3.6 Flash
94.1%
xGrok 4.6
94.0%
OGPT-5.5
94.0%
OGPT-5.5 Pro
93.9%
AnthropicClaude Opus 5
93.9%
OGPT-5.6 Sol
93.5%
xGrok 4.5
93.4%
OGPT-5.6 Terra
93.3%
OGPT-5.4
93.3%
Moonshot AIKimi K3
93.1%
GoogleGemini 3.5 Flash
92.8%
xGrok 4.7
92.7%

Bars show accuracy (%) and start at 0%.

View table (top 30)
RankModelDeveloperRelease dateScore
1#1OGPT-6 Astra(max)OpenAI2026-09-0395.8%
2#2AnthropicClaude Sonnet 5.5(max)Anthropic2026-09-2895.6%
3#3OGPT-6.1 Sol(max)OpenAI2026-09-2995.4%
4#4GoogleGemini 3.8 Flash(high)Google2026-09-0295.4%
5#5GoogleGemini 3.7 Flash(high)Google2026-08-1394.8%
6#6OGPT-5.4 Pro(xhigh)OpenAI2026-03-0594.6%
7#7GoogleGemini 3.1 Pro PreviewGoogle2026-02-1994.4%
8#8OGPT-6 Sol(max)OpenAI2026-09-2294.3%
9#9GoogleGemini 3.6 FlashGoogle2026-07-2194.1%
10#10xGrok 4.6(high)xAI2026-08-1294.0%
11#11OGPT-5.5(xhigh)OpenAI2026-04-2394.0%
12#12OGPT-5.5 Pro(xhigh)OpenAI2026-04-2393.9%
13#13AnthropicClaude Opus 5(max)Anthropic2026-07-2493.9%
14#14OGPT-5.6 Sol(max)OpenAI2026-07-0993.5%
15#15xGrok 4.5(high)xAI2026-07-0893.4%
16#16OGPT-5.6 Terra(max)OpenAI2026-07-0993.3%
17#17OGPT-5.4(xhigh)OpenAI2026-03-0593.3%
18#18Moonshot AIKimi K3(max)Moonshot AI2026-07-1693.1%
19#19GoogleGemini 3.5 FlashGoogle2026-05-1992.8%
20#20xGrok 4.7(xhigh)xAI2026-09-2192.7%
21#21AlibabaQwen3.8 Max(xhigh)Alibaba2026-08-0292.7%
22#22GoogleGemini 3 Pro PreviewGoogle2025-11-1892.6%
23#23AlibabaQwen3.8 Max (0902)(xhigh)Alibaba2026-09-0192.3%
24#24ZGLM-5.2Z.ai2026-06-1691.9%
25#25DeepSeekDeepSeek V4 Pro 0813(max)DeepSeek2026-08-1391.7%
26#26OGPT-5.6 Luna(max)OpenAI2026-07-0991.6%
27#27OGPT-5.2(xhigh)OpenAI2025-12-1191.4%
28#28AnthropicClaude Opus 4.8Anthropic2026-05-2891.0%
29#29DeepSeekDeepSeek V4 Flash 0731(max)DeepSeek2026-07-3191.0%
30#30ZGLM-5.3(max)Z.ai2026-08-1490.9%

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings