Valumigo

AI benchmark rankings

We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.

Latest models at a glance

Models released in the past 120 days, sorted by release date, with scores from both public sources in one row. Test scores show accuracy (%).

ModelRelease dateOverall indexGPQA DiamondFrontierMathHLEARC-AGI-2SWE-bench VerifiedLMArena overall rankLMArena English rank
OGPT-6.1 SolOpenAI2026-09-29–95.493.7–––#60#57
AnthropicClaude Sonnet 5.5Anthropic2026-09-28165.295.688.8–––#34#16
AnthropicClaude Opus 5.5Anthropic2026-09-22167.4–91.2–92.5–#2#2
OGPT-6 SolOpenAI2026-09-22–94.389.8–89.6–#158#162
OGPT-6 LunaOpenAI2026-09-22––78.9–59.3–#163#171
DeepSeekDeepSeek V4.1 FlashDeepSeek2026-09-09155–––––#37#30
OGPT-6 AstraOpenAI2026-09-03166.595.893.754.895–#71#76
GoogleGemini 3.8 FlashGoogle2026-09-02156.995.468.444.5––#8#10
MetaMuse Spark 1.3Meta2026-09-02156.9–74.4–––#12#12
AnthropicClaude Fable 5.1Anthropic2026-09-01164.8–90.246.590–#3#7
AlibabaQwen3.8 Max (0902)Alibaba2026-09-01155.292.365.6–––Not yet availableNot yet available
ZGLM-5.3-FlashZ.ai2026-08-20151.9–––––#30#34
ZGLM-5.3Z.ai2026-08-14155.890.968.8–––#26#19
AlibabaQwen 3.8 27BAlibaba2026-08-14149.4–––––#74#81
GoogleGemini 3.7 FlashGoogle2026-08-13157.494.871.6–84.6–#13#17
DeepSeekDeepSeek V4 Pro 0813DeepSeek2026-08-13155.491.7––61.3–Not yet availableNot yet available

New models need enough votes to appear on leaderboards, especially language-specific rankings. They may show as 'Not yet available' shortly after release.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

English LMArena

Published 2026-10-02 · Retrieved 2026-10-04

1470150015301560
Googlegemini-4-argon-high
1539
Anthropicclaude-opus-5.5-high
1513
Anthropicclaude-opus-4-6-high
1512
Anthropicclaude-opus-4-6
1506
Anthropicclaude-opus-5-max
1505
Xiaomimimo-v2.6-pro
1504
Anthropicclaude-fable-5.1-max
1504
Anthropicclaude-opus-5-high
1502
Anthropicclaude-fable-5-high
1500
Googlegemini-3.8-flash-high
1498
Anthropicclaude-opus-4-7-high
1496
Metamuse-spark-1.3-max
1492
Alibabaqwen3.8-max
1492
Metamuse-spark-1.2 (xHigh)
1489
Anthropicclaude-opus-4-7
1488
Anthropicclaude-sonnet-5.5-xhigh
1487
Googlegemini-3.7-flash-high
1484
Moonshot AIkimi-k3-max
1482
Zglm-5.3-max
1482
Googlegemini-3.6-flash-high
1482

Dots show scores; horizontal lines show 95% confidence intervals. Overlapping intervals suggest similar performance.

View table (top 50)
RankModelDeveloperScore95% confidence intervalVotes
1Googlegemini-4-argon-highGoogle15391525–15531,928
2Anthropicclaude-opus-5.5-highAnthropic15131498–15271,719
3Anthropicclaude-opus-4-6-highAnthropic15121508–151734,324
4Anthropicclaude-opus-4-6Anthropic15061502–151136,805
5Anthropicclaude-opus-5-maxAnthropic15051498–151111,702
6Xiaomimimo-v2.6-proXiaomi15041489–15191,487
7Anthropicclaude-fable-5.1-maxAnthropic15041495–15124,793
8Anthropicclaude-opus-5-highAnthropic15021496–150724,266
9Anthropicclaude-fable-5-highAnthropic15001494–150516,474
10Googlegemini-3.8-flash-highGoogle14981491–150510,363
11Anthropicclaude-opus-4-7-highAnthropic14961491–150130,082
12Metamuse-spark-1.3-maxMeta14921484–15014,928
13Alibabaqwen3.8-maxAlibaba14921485–14999,003
14Metamuse-spark-1.2 (xHigh)Meta14891475–15021,805
15Anthropicclaude-opus-4-7Anthropic14881483–149330,422
16Anthropicclaude-sonnet-5.5-xhighAnthropic14871470–15041,196
17Googlegemini-3.7-flash-highGoogle14841477–14918,659
18Moonshot AIkimi-k3-maxMoonshot AI14821476–148910,786
19Zglm-5.3-maxZ.ai14821474–14907,223
20Googlegemini-3.6-flash-highGoogle14821476–148815,114
21Metamuse-sparkMeta14801473–14886,673
22Googlegemini-3.5-flash-highGoogle14801475–148620,509
23Metamuse-spark-1.1Meta14791473–148515,100
24Googlegemini-3.1-pro-previewGoogle14791475–148353,727
25Xiaomimimo-v2.6-flashXiaomi14781466–14912,266
26Baiduernie-5.1Baidu14781472–148417,926
27Googlegemini-3-proGoogle14781472–148316,887
28Xiaomimimo-v2.5-proXiaomi14781473–148229,857
29Googlegemini-3.5-flash-mediumGoogle14761471–148119,998
30DeepSeekdeepseek-v4.1-flash-maxDeepSeek14761466–14863,237
31Zglm-5.1Z.ai14761471–148025,210
32Alibabaqwen3.5-max-previewAlibaba14751469–148210,143
33Zglm-5.2-maxZ.ai14741469–148018,859
34Zglm-5.3-flashZ.ai14731466–14808,546
35Anthropicclaude-sonnet-4-6Anthropic14721467–147731,850
36Alibabaqwen3.7-max-previewAlibaba14701457–14841,884
37Ogpt-5.5-highOpenAI14701465–147531,442
38Ogpt-5.5OpenAI14691464–147432,123
39Anthropicclaude-opus-4-8-highAnthropic14681463–147327,922
40Googlegemini-3-flashGoogle14671461–147312,474
41Ogpt-5.4-highOpenAI14671462–147229,222
42SStep 5 PreviewStepfun14611442–1479961
43Anthropicclaude-opus-4-5-20251101Anthropic14611456–146531,143
44Anthropicclaude-opus-4-8Anthropic14601455–146528,640
45Moonshot AIkimi-k2.6Moonshot AI14601455–146618,104
46DeepSeekdeepseek-v4-proDeepSeek14591454–146425,209
47Alibabaqwen3.7-plusAlibaba14581453–146417,608
48Zglm-5Z.ai14581452–146412,877
49Googlegemini-2.5-proGoogle14581455–146153,818
50Aamazon-nova-experimental-chat-26-02-10Amazon14581442–14731,310

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks