Valumigo

LMArena · Agent task success

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: This ranking isolates the 'task success confirmation' signal from Agent Arena. It uses results that users explicitly mark as successful or unsuccessful after an agent completes a task.

How to read the score: Scores use IPS(τ̂), with 0 as the baseline. Higher scores indicate better user-reported success outcomes relative to the baseline.

Caveats: Only tasks explicitly marked by users are included; unmarked tasks are excluded. The order may differ from the overall ranking.

Insights from this ranking

  • The 95% confidence intervals for 1st-place Claude Fable 5.1 (Max) and 2nd-place GPT 6.1 Sol (Max) overlap. This aggregation alone does not clearly establish their order.
  • 5 other models have confidence intervals that overlap with the 1st-place model's. Interpret small ranking differences cautiously alongside vote counts.
  • Anthropic has the most models among the top 10, with 5.
  • The top-ranked model has 7,868 observations, and this leaderboard includes 50 models.

Agent task success LMArena

Published 2026-10-02 · Retrieved 2026-10-04

00.050.10.150.20.25
AnthropicClaude Fable 5.1 (Max)
0.176
OGPT 6.1 Sol (Max)
0.155
GoogleGemini 4 Argon (High)
0.154
AnthropicClaude Opus 5.5 (High)
0.141
AnthropicClaude Sonnet 5.5 (Max)
0.139
OGPT 6 Astra (Max)
0.132
AnthropicClaude Opus 5 (Max)
0.095
DeepSeekDeepseek V4.1 Flash (Max)
0.084
Moonshot AIKimi K3 (Max)
0.076
OGPT 6 Sol (Max)
0.071
GoogleGemini 3.8 Flash (High)
0.069
AnthropicClaude Fable 5 (High)
0.061
THy4 preview
0.058
XiaomiMiMo V2.6 Flash
0.056
xGrok 4.7 (xHigh)
0.055
ZGLM 5.3 (Max)
0.053
OGPT 5.6 Sol (xHigh)
0.05
MetaMuse Spark 1.3 (Max)
0.05
ZGLM 5.2 (Max)
0.048
AlibabaQwen3.8 Max
0.047

Scores use IPS(τ̂), with 0 as the baseline and higher scores indicating better performance. Lines show 95% confidence intervals. The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings.

Top ranking by company

  • AnthropicAnthropicClaude Fable 5.1 (Max)#1
  • OOpenAIGPT 6.1 Sol (Max)#2
  • GoogleGoogleGemini 4 Argon (High)#3
  • DeepSeekDeepSeekDeepseek V4.1 Flash (Max)#8
  • Moonshot AIMoonshot AIKimi K3 (Max)#9
  • TTencentHy4 preview#14
  • XiaomiXiaomiMiMo V2.6 Flash#15
  • xxAIGrok 4.7 (xHigh)#16
  • ZZ.aiGLM 5.3 (Max)#17
View table (top 50)
RankModelDeveloperScore95% confidence intervalObservations
1#1AnthropicClaude Fable 5.1 (Max)Anthropic0.1760.148–0.2057,868
2#2OGPT 6.1 Sol (Max)OpenAI0.1550.114–0.1962,270
3#3GoogleGemini 4 Argon (High)Google0.1540.125–0.1844,422
4#4AnthropicClaude Opus 5.5 (High)Anthropic0.1410.101–0.1823,425
5#5AnthropicClaude Sonnet 5.5 (Max)Anthropic0.1390.086–0.1931,195
6#6OGPT 6 Astra (Max)OpenAI0.1320.095–0.1695,954
7#7AnthropicClaude Opus 5 (Max)Anthropic0.0950.064–0.12615,453
8#8DeepSeekDeepseek V4.1 Flash (Max)DeepSeek0.0840.074–0.09560,410
9#9Moonshot AIKimi K3 (Max)Moonshot AI0.0760.064–0.08899,685
10#10AnthropicClaude Opus 5 (High)Anthropic0.0760.046–0.10620,843
11#11OGPT 6 Sol (Max)OpenAI0.0710.028–0.1143,312
12#12GoogleGemini 3.8 Flash (High)Google0.0690.05–0.08722,497
13#13AnthropicClaude Fable 5 (High)Anthropic0.0610.035–0.08733,005
14#14THy4 previewTencent0.0580.042–0.07524,238
15#15XiaomiMiMo V2.6 FlashXiaomi0.0560.03–0.0827,967
16#16xGrok 4.7 (xHigh)xAI0.0550.021–0.0895,906
17#17ZGLM 5.3 (Max)Z.ai0.0530.037–0.06862,886
18#18OGPT 5.6 Sol (xHigh)OpenAI0.050.024–0.07728,118
19#19MetaMuse Spark 1.3 (Max)Meta0.050.037–0.06445,563
20#20ZGLM 5.2 (Max)Z.ai0.0480.032–0.06355,555
21#21AlibabaQwen3.8 MaxAlibaba0.0470.033–0.06135,416
22#22SStep 5 PreviewStepFun0.0450.012–0.0784,266
23#23ZGLM 5.3 FlashZ.ai0.0440.032–0.05560,370
24#24AnthropicClaude Opus 4.8 (High)Anthropic0.0420.016–0.06930,039
25#25AlibabaQwen3.8 Flash NextAlibaba0.0360.021–0.05138,898
26#26xGrok 4.5xAI0.013-0.01–0.03729,109
27#27AnthropicClaude Sonnet 5 (High)Anthropic0.012-0.023–0.04722,551
28#28GoogleGemini 3.7 Flash (High)Google0.012-0.006–0.02948,036
29#29AlibabaQwen 3.8 27BAlibaba0.01-0.005–0.02436,984
30#30OGPT 6 Luna (Max)OpenAI-0.001-0.025–0.02314,416
31#31xGrok 4.6 (xHigh)xAI-0.006-0.029–0.01820,463
32#32OGPT 5.5 (xHigh)OpenAI-0.009-0.031–0.01443,172
33#33DeepSeekDeepSeek V4 Pro (High) (0813)DeepSeek-0.017-0.057–0.0233,790
34#34OGPT 5.4 (High)OpenAI-0.025-0.046–-0.00457,450
35#35OGPT 5.6 Luna (xHigh)OpenAI-0.033-0.049–-0.01733,799
36#36OGPT 5.6 Terra (xHigh)OpenAI-0.039-0.065–-0.01318,613
37#37MetaMuse Spark 1.1Meta-0.043-0.054–-0.03195,130
38#38OGPT 5.5OpenAI-0.047-0.068–-0.02664,950
39#39MetaMuse Spark 1.2 (xHigh)Meta-0.054-0.072–-0.03633,552
40#40GoogleGemini 3.6 Flash (High)Google-0.057-0.083–-0.03118,029
41#41AlibabaQwen3.7 MaxAlibaba-0.091-0.115–-0.06630,156
42#42AlibabaQwen3.7 PlusAlibaba-0.093-0.126–-0.0615,130
43#43XiaomiMimo V2.5 ProXiaomi-0.096-0.119–-0.07329,517
44#44GoogleGemini 3.1 Pro PreviewGoogle-0.099-0.121–-0.07869,866
45#45THy3Tencent-0.112-0.14–-0.08414,238
46#46MiniMaxMinimax M3MiniMax-0.137-0.162–-0.11228,148
47#47Mistral AIMistral Medium 3.5Mistral AI-0.145-0.183–-0.1086,629
48#48TInkling SmallThinky-0.201-0.244–-0.1587,271
49#49USolar Pro 4Upstage-0.205-0.259–-0.1524,788
50#50TInklingThinky-0.207-0.237–-0.17724,644

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings