Valumigo

LMArena · Overall agent performance

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: This ranking is based on records from LMArena's Agent Arena, where people assign real tasks, such as software development and financial analysis, to models in agent mode. Agents directly use tools such as bash, web search, and file operations.

How to read the score: The score is a causal estimate called IPS(τ̂), with 0 as the baseline. Higher scores indicate that using the model led to better outcomes. It combines five signals: user confirmation of success, praise or complaints, incorporation of correction requests, recovery from bash errors, and avoidance of tool hallucinations.

Caveats: The score scale differs from chat rankings, which use Bradley–Terry, so they cannot be compared directly. Models with fewer observations have wider confidence intervals.

Insights from this ranking

  • The 95% confidence intervals for 1st-place Claude Fable 5.1 (Max) and 2nd-place Claude Opus 5.5 (High) overlap. This aggregation alone does not clearly establish their order.
  • 4 other models have confidence intervals that overlap with the 1st-place model's. Interpret small ranking differences cautiously alongside vote counts.
  • Anthropic has the most models among the top 10, with 6.
  • The top-ranked model has 1,683,381 observations, and this leaderboard includes 50 models.

Overall agent performance LMArena

Published 2026-10-02 · Retrieved 2026-10-04

00.040.080.120.160.2
AnthropicClaude Fable 5.1 (Max)
0.143
AnthropicClaude Opus 5.5 (High)
0.138
AnthropicClaude Sonnet 5.5 (Max)
0.125
OGPT 6 Astra (Max)
0.123
OGPT 6.1 Sol (Max)
0.112
OGPT 6 Sol (Max)
0.097
AnthropicClaude Opus 5 (High)
0.087
AnthropicClaude Fable 5 (High)
0.082
GoogleGemini 4 Argon (High)
0.076
AnthropicClaude Opus 4.8 (High)
0.066
OGPT 5.6 Sol (xHigh)
0.065
AnthropicClaude Sonnet 5 (High)
0.044
Moonshot AIKimi K3 (Max)
0.042
OGPT 5.5 (xHigh)
0.042
xGrok 4.7 (xHigh)
0.04
DeepSeekDeepseek V4.1 Flash (Max)
0.04
MetaMuse Spark 1.3 (Max)
0.04
THy4 preview
0.04
ZGLM 5.2 (Max)
0.035
GoogleGemini 3.8 Flash (High)
0.03

Scores use IPS(τ̂), with 0 as the baseline and higher scores indicating better performance. Lines show 95% confidence intervals. The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings.

Top ranking by company

  • AnthropicAnthropicClaude Fable 5.1 (Max)#1
  • OOpenAIGPT 6 Astra (Max)#4
  • GoogleGoogleGemini 4 Argon (High)#10
  • Moonshot AIMoonshot AIKimi K3 (Max)#14
  • xxAIGrok 4.7 (xHigh)#16
  • DeepSeekDeepSeekDeepseek V4.1 Flash (Max)#17
  • MetaMetaMuse Spark 1.3 (Max)#18
  • TTencentHy4 preview#19
  • ZZ.aiGLM 5.2 (Max)#20
View table (top 50)
RankModelDeveloperScore95% confidence intervalObservations
1#1AnthropicClaude Fable 5.1 (Max)Anthropic0.1430.124–0.1621,683,381
2#2AnthropicClaude Opus 5.5 (High)Anthropic0.1380.116–0.16754,031
3#3AnthropicClaude Sonnet 5.5 (Max)Anthropic0.1250.094–0.156483,211
4#4OGPT 6 Astra (Max)OpenAI0.1230.1–0.145888,641
5#5OGPT 6.1 Sol (Max)OpenAI0.1120.085–0.14379,372
6#6OGPT 6 Sol (Max)OpenAI0.0970.073–0.121994,367
7#7AnthropicClaude Opus 5 (High)Anthropic0.0870.073–0.1013,531,760
8#8AnthropicClaude Fable 5 (High)Anthropic0.0820.07–0.0942,603,067
9#9AnthropicClaude Opus 5 (Max)Anthropic0.0790.064–0.0943,006,257
10#10GoogleGemini 4 Argon (High)Google0.0760.056–0.095587,353
11#11AnthropicClaude Opus 4.8 (High)Anthropic0.0660.053–0.082,431,185
12#12OGPT 5.6 Sol (xHigh)OpenAI0.0650.052–0.0774,348,715
13#13AnthropicClaude Sonnet 5 (High)Anthropic0.0440.028–0.063,612,898
14#14Moonshot AIKimi K3 (Max)Moonshot AI0.0420.036–0.04712,305,667
15#15OGPT 5.5 (xHigh)OpenAI0.0420.031–0.0523,060,353
16#16xGrok 4.7 (xHigh)xAI0.040.024–0.0561,656,641
17#17DeepSeekDeepseek V4.1 Flash (Max)DeepSeek0.040.036–0.04516,932,092
18#18MetaMuse Spark 1.3 (Max)Meta0.040.034–0.0466,139,426
19#19THy4 previewTencent0.040.032–0.0475,108,255
20#20ZGLM 5.2 (Max)Z.ai0.0350.028–0.0426,684,620
21#21GoogleGemini 3.8 Flash (High)Google0.030.021–0.0384,279,699
22#22AlibabaQwen3.8 MaxAlibaba0.0250.019–0.0314,577,924
23#23ZGLM 5.3 (Max)Z.ai0.0240.017–0.039,155,315
24#24OGPT 6 Luna (Max)OpenAI0.0140.003–0.0243,443,780
25#25xGrok 4.6 (xHigh)xAI0.0130.003–0.0233,131,725
26#26xGrok 4.5xAI0.0120.002–0.0222,418,629
27#27DeepSeekDeepSeek V4 Pro (High) (0813)DeepSeek0.012-0.005–0.028797,284
28#28OGPT 5.4 (High)OpenAI0.0110.001–0.024,332,904
29#29OGPT 5.5OpenAI0.010.001–0.022,659,460
30#30SStep 5 PreviewStepFun0.005-0.01–0.021,125,598
31#31OGPT 5.6 Terra (xHigh)OpenAI0.004-0.008–0.0161,686,712
32#32ZGLM 5.3 FlashZ.ai-0.004-0.009–09,533,773
33#33XiaomiMiMo V2.6 FlashXiaomi-0.006-0.019–0.0071,562,757
34#34AlibabaQwen3.8 Flash NextAlibaba-0.007-0.014–-0.00111,249,221
35#35OGPT 5.6 Luna (xHigh)OpenAI-0.012-0.018–-0.0054,028,029
36#36GoogleGemini 3.7 Flash (High)Google-0.015-0.022–-0.0075,048,920
37#37AlibabaQwen 3.8 27BAlibaba-0.017-0.023–-0.0116,972,838
38#38MetaMuse Spark 1.2 (xHigh)Meta-0.033-0.039–-0.0262,884,900
39#39MetaMuse Spark 1.1Meta-0.049-0.053–-0.0447,612,456
40#40AlibabaQwen3.7 MaxAlibaba-0.051-0.061–-0.0422,207,521
41#41THy3Tencent-0.054-0.066–-0.0421,171,568
42#42MiniMaxMinimax M3MiniMax-0.069-0.078–-0.063,338,791
43#43AlibabaQwen3.7 PlusAlibaba-0.072-0.086–-0.0571,144,063
44#44XiaomiMimo V2.5 ProXiaomi-0.075-0.084–-0.0662,309,389
45#45GoogleGemini 3.6 Flash (High)Google-0.077-0.087–-0.0661,523,777
46#46GoogleGemini 3.1 Pro PreviewGoogle-0.077-0.087–-0.0673,446,977
47#47TInkling SmallThinky-0.103-0.117–-0.089503,653
48#48TInklingThinky-0.109-0.119–-0.0981,798,990
49#49Mistral AIMistral Medium 3.5Mistral AI-0.124-0.139–-0.109437,621
50#50USolar Pro 4Upstage-0.159-0.178–-0.141427,705

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings