LMArena · Agent task success
We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.
How to read this leaderboard
What it measures: This ranking isolates the 'task success confirmation' signal from Agent Arena. It uses results that users explicitly mark as successful or unsuccessful after an agent completes a task.
How to read the score: Scores use IPS(τ̂), with 0 as the baseline. Higher scores indicate better user-reported success outcomes relative to the baseline.
Caveats: Only tasks explicitly marked by users are included; unmarked tasks are excluded. The order may differ from the overall ranking.
Insights from this ranking
- The 95% confidence intervals for 1st-place Claude Fable 5.1 (Max) and 2nd-place GPT 6.1 Sol (Max) overlap. This aggregation alone does not clearly establish their order.
- 5 other models have confidence intervals that overlap with the 1st-place model's. Interpret small ranking differences cautiously alongside vote counts.
- Anthropic has the most models among the top 10, with 5.
- The top-ranked model has 7,868 observations, and this leaderboard includes 50 models.
Agent task success LMArena
Published 2026-10-02 · Retrieved 2026-10-04
Scores use IPS(τ̂), with 0 as the baseline and higher scores indicating better performance. Lines show 95% confidence intervals. The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings.
Top ranking by company
- AnthropicClaude Fable 5.1 (Max)#1
- OOpenAIGPT 6.1 Sol (Max)#2
- GoogleGemini 4 Argon (High)#3
- DeepSeekDeepseek V4.1 Flash (Max)#8
- Moonshot AIKimi K3 (Max)#9
- TTencentHy4 preview#14
- XiaomiMiMo V2.6 Flash#15
- xxAIGrok 4.7 (xHigh)#16
- ZZ.aiGLM 5.3 (Max)#17
View table (top 50)
| Rank | Model | Developer | Score | 95% confidence interval | Observations |
|---|---|---|---|---|---|
| 1 | #1Claude Fable 5.1 (Max) | Anthropic | 0.176 | 0.148–0.205 | 7,868 |
| 2 | #2OGPT 6.1 Sol (Max) | OpenAI | 0.155 | 0.114–0.196 | 2,270 |
| 3 | #3Gemini 4 Argon (High) | 0.154 | 0.125–0.184 | 4,422 | |
| 4 | #4Claude Opus 5.5 (High) | Anthropic | 0.141 | 0.101–0.182 | 3,425 |
| 5 | #5Claude Sonnet 5.5 (Max) | Anthropic | 0.139 | 0.086–0.193 | 1,195 |
| 6 | #6OGPT 6 Astra (Max) | OpenAI | 0.132 | 0.095–0.169 | 5,954 |
| 7 | #7Claude Opus 5 (Max) | Anthropic | 0.095 | 0.064–0.126 | 15,453 |
| 8 | #8Deepseek V4.1 Flash (Max) | DeepSeek | 0.084 | 0.074–0.095 | 60,410 |
| 9 | #9Kimi K3 (Max) | Moonshot AI | 0.076 | 0.064–0.088 | 99,685 |
| 10 | #10Claude Opus 5 (High) | Anthropic | 0.076 | 0.046–0.106 | 20,843 |
| 11 | #11OGPT 6 Sol (Max) | OpenAI | 0.071 | 0.028–0.114 | 3,312 |
| 12 | #12Gemini 3.8 Flash (High) | 0.069 | 0.05–0.087 | 22,497 | |
| 13 | #13Claude Fable 5 (High) | Anthropic | 0.061 | 0.035–0.087 | 33,005 |
| 14 | #14THy4 preview | Tencent | 0.058 | 0.042–0.075 | 24,238 |
| 15 | #15MiMo V2.6 Flash | Xiaomi | 0.056 | 0.03–0.082 | 7,967 |
| 16 | #16xGrok 4.7 (xHigh) | xAI | 0.055 | 0.021–0.089 | 5,906 |
| 17 | #17ZGLM 5.3 (Max) | Z.ai | 0.053 | 0.037–0.068 | 62,886 |
| 18 | #18OGPT 5.6 Sol (xHigh) | OpenAI | 0.05 | 0.024–0.077 | 28,118 |
| 19 | #19Muse Spark 1.3 (Max) | Meta | 0.05 | 0.037–0.064 | 45,563 |
| 20 | #20ZGLM 5.2 (Max) | Z.ai | 0.048 | 0.032–0.063 | 55,555 |
| 21 | #21Qwen3.8 Max | Alibaba | 0.047 | 0.033–0.061 | 35,416 |
| 22 | #22SStep 5 Preview | StepFun | 0.045 | 0.012–0.078 | 4,266 |
| 23 | #23ZGLM 5.3 Flash | Z.ai | 0.044 | 0.032–0.055 | 60,370 |
| 24 | #24Claude Opus 4.8 (High) | Anthropic | 0.042 | 0.016–0.069 | 30,039 |
| 25 | #25Qwen3.8 Flash Next | Alibaba | 0.036 | 0.021–0.051 | 38,898 |
| 26 | #26xGrok 4.5 | xAI | 0.013 | -0.01–0.037 | 29,109 |
| 27 | #27Claude Sonnet 5 (High) | Anthropic | 0.012 | -0.023–0.047 | 22,551 |
| 28 | #28Gemini 3.7 Flash (High) | 0.012 | -0.006–0.029 | 48,036 | |
| 29 | #29Qwen 3.8 27B | Alibaba | 0.01 | -0.005–0.024 | 36,984 |
| 30 | #30OGPT 6 Luna (Max) | OpenAI | -0.001 | -0.025–0.023 | 14,416 |
| 31 | #31xGrok 4.6 (xHigh) | xAI | -0.006 | -0.029–0.018 | 20,463 |
| 32 | #32OGPT 5.5 (xHigh) | OpenAI | -0.009 | -0.031–0.014 | 43,172 |
| 33 | #33DeepSeek V4 Pro (High) (0813) | DeepSeek | -0.017 | -0.057–0.023 | 3,790 |
| 34 | #34OGPT 5.4 (High) | OpenAI | -0.025 | -0.046–-0.004 | 57,450 |
| 35 | #35OGPT 5.6 Luna (xHigh) | OpenAI | -0.033 | -0.049–-0.017 | 33,799 |
| 36 | #36OGPT 5.6 Terra (xHigh) | OpenAI | -0.039 | -0.065–-0.013 | 18,613 |
| 37 | #37Muse Spark 1.1 | Meta | -0.043 | -0.054–-0.031 | 95,130 |
| 38 | #38OGPT 5.5 | OpenAI | -0.047 | -0.068–-0.026 | 64,950 |
| 39 | #39Muse Spark 1.2 (xHigh) | Meta | -0.054 | -0.072–-0.036 | 33,552 |
| 40 | #40Gemini 3.6 Flash (High) | -0.057 | -0.083–-0.031 | 18,029 | |
| 41 | #41Qwen3.7 Max | Alibaba | -0.091 | -0.115–-0.066 | 30,156 |
| 42 | #42Qwen3.7 Plus | Alibaba | -0.093 | -0.126–-0.06 | 15,130 |
| 43 | #43Mimo V2.5 Pro | Xiaomi | -0.096 | -0.119–-0.073 | 29,517 |
| 44 | #44Gemini 3.1 Pro Preview | -0.099 | -0.121–-0.078 | 69,866 | |
| 45 | #45THy3 | Tencent | -0.112 | -0.14–-0.084 | 14,238 |
| 46 | #46Minimax M3 | MiniMax | -0.137 | -0.162–-0.112 | 28,148 |
| 47 | #47Mistral Medium 3.5 | Mistral AI | -0.145 | -0.183–-0.108 | 6,629 |
| 48 | #48TInkling Small | Thinky | -0.201 | -0.244–-0.158 | 7,271 |
| 49 | #49USolar Pro 4 | Upstage | -0.205 | -0.259–-0.152 | 4,788 |
| 50 | #50TInkling | Thinky | -0.207 | -0.237–-0.177 | 24,644 |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings