Valumigo

OpenRouter · Graduate-level science questions (GPQA Diamond)

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: Accuracy on GPQA Diamond (graduate-level multiple-choice questions in biology, chemistry, and physics), evaluated independently by OpenRouter. Average cost per task is also published.

How to read the score: Higher accuracy (%) is better. 'Average cost per task' is the average amount spent solving one question at OpenRouter rates, allowing you to compare cost against performance.

Caveats: Evaluation conditions differ from those used for Epoch AI's GPQA scores, so the figures may differ. Costs are based on OpenRouter rates.

Insights from this ranking

  • The highest score is 95.6%, achieved by Gemini 3.8 Flash.
  • The gap between 1st and 5th place is 1.2%p.
  • OpenAI has the most models among the top 10, with 5.
  • Among the top 10, Jev Router has the lowest average cost per task ($0.0159, score 93.9%).

Graduate-level science questions (GPQA Diamond) OpenRouter

Published 2026-10-04 · Retrieved 2026-10-05

GoogleGemini 3.8 Flash
95.6%
OGPT-6 Astra Pro
95.5%
SFugu Ultra
94.6%
GoogleGemini 3.1 Pro Preview
94.5%
OGPT-6 Astra
94.4%
OGPT-6.1 Sol
94.4%
TJev Router
93.9%
OGPT-5.6 Sol Pro
93.9%
GoogleGemini 3.7 Flash
93.9%
OGPT-5.5
93.7%
SFugu Ultra V2
93.3%
xGrok 4.6
92.8%
GoogleGemini 3.6 Flash
92.6%
GoogleGemini 3.5 Flash
92.6%
UPareto
92.4%
OGPT-5.6 Sol
92.1%
AnthropicClaude Sonnet 5.5
92.1%
OGPT-6 Sol
91.9%
Moonshot AIKimi K3
91.9%
NVIDIASwitchyard
91.4%

The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings.

View table (top 60)
RankModelDeveloperScoreAverage cost per task
1#1GoogleGemini 3.8 FlashGoogle95.6%$0.0682
2#2OGPT-6 Astra ProOpenAI95.5%$0.3326
3#3SFugu UltraSakana94.6%$1.1487
4#4GoogleGemini 3.1 Pro PreviewGoogle94.5%$0.202
5#5OGPT-6 AstraOpenAI94.4%$0.1133
6#6OGPT-6.1 SolOpenAI94.4%$0.0214
7#7TJev RouterTypesafe93.9%$0.0159
8#8OGPT-5.6 Sol ProOpenAI93.9%$0.2979
9#9GoogleGemini 3.7 FlashGoogle93.9%$0.0293
10#10OGPT-5.5OpenAI93.7%$0.3093
11#11SFugu Ultra V2Sakana93.3%$0.5026
12#12xGrok 4.6xAI92.8%$0.1233
13#13GoogleGemini 3.6 FlashGoogle92.6%$0.0596
14#14GoogleGemini 3.5 FlashGoogle92.6%$0.1413
15#15UParetoUnbiased92.4%$0.0338
16#16OGPT-5.6 SolOpenAI92.1%$0.0604
17#17AnthropicClaude Sonnet 5.5Anthropic92.1%$0.0253
18#18OGPT-6 SolOpenAI91.9%$0.0253
19#19Moonshot AIKimi K3Moonshot AI91.9%$0.1323
20#20NVIDIASwitchyardNVIDIA91.4%$0.0179
21#21ZGLM 5.3 FlashZ.ai90.9%$0.0054
22#22AlibabaQwen3.7 MaxAlibaba90.9%$0.2222
23#23MiniMaxMiniMax M3MiniMax90.3%$0.0348
24#24OGPT-5.4OpenAI90.3%$0.1467
25#25AnthropicClaude Opus 5.5Anthropic90.2%$0.0542
26#26OGPT-5.6 Luna ProOpenAI90%$0.0412
27#27AnthropicClaude Fable 5.1Anthropic89.9%$0.1305
28#28DeepSeekDeepSeek V4 Pro 0813DeepSeek89.8%$0.1143
29#29AnthropicClaude Opus 4.8Anthropic89.7%$0.2016
30#30AnthropicClaude Opus 4.7Anthropic89.4%$0.2239
31#31THy4 previewTencent89.1%$0.1253
32#32ANova Micro 1.0Amazon89%$0.2165
33#33DeepSeekDeepSeek V4.1 FlashDeepSeek88.9%$0.0244
34#34OGPT-5.6 TerraOpenAI88.8%$0.0344
35#35XiaomiMiMo-V2.6-ProXiaomi88.6%$0.0319
36#36SFugu MaxSakana88.6%$0.0756
37#37AlibabaQwen3.8 FlashAlibaba88.6%$0.0215
38#38GoogleGemini 3 Flash PreviewGoogle88.4%$0.1318
39#39OGPT-6 Luna ProOpenAI88.4%$0.0083
40#40DeepSeekDeepSeek V4 Flash Vision ExpDeepSeek88%$0.0228
41#41OAuto RouterOpenrouter88%$0.0892
42#42OGPT-5.2OpenAI87.9%$0.1333
43#43AlibabaQwen3.7 PlusAlibaba87.9%$0.044
44#44OGPT-5.6 LunaOpenAI87.9%$0.0079
45#45AnthropicClaude Opus 5Anthropic87.5%$0.1306
46#46OGPT-6 LunaOpenAI87.4%$0.003
47#47Moonshot AIKimi K2 ThinkingMoonshot AI87.3%$0.1062
48#48AlibabaQwen3.8 2.4T A95BAlibaba86.9%$0.0824
49#49TInkling SmallThinkingmachines86.6%$0.0253
50#50OGPT-5.1OpenAI86.4%$0.2213
51#51DeepSeekDeepSeek V4 Flash 0731DeepSeek86.3%$0.008
52#52MiniMaxMiniMax M2.1MiniMax86.2%$0.0403
53#53AnthropicClaude Opus 4.5Anthropic86.2%$0.8454
54#54DeepSeekDeepSeek V4 Pro 0423DeepSeek86.2%$0.0544
55#55AnthropicClaude Opus 4.6Anthropic86%$0.7469
56#56DeepSeekDeepSeek V4 Flash 0423DeepSeek85.8%$0.0062
57#57ZGLM 5.3 FlashZ.ai85.7%$0.0123
58#58OGPT-5OpenAI85.7%$0.2099
59#59AnthropicClaude Fable 5Anthropic85.5%$0.1754
60#60ZGLM 5.3Z.ai85.4%$0.0925

Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). Licensed under CC BY 4.0.

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings). Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). Source: Design Arena (www.designarena.ai) via OpenRouter (openrouter.ai/rankings). Licensed under CC BY 4.0. Source: OpenRouter (openrouter.ai/rankings), as of 2026-10-05. Licensed under CC BY 4.0. https://openrouter.ai/rankings