LMArena · Overall agent performance
We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.
How to read this leaderboard
What it measures: This ranking is based on records from LMArena's Agent Arena, where people assign real tasks, such as software development and financial analysis, to models in agent mode. Agents directly use tools such as bash, web search, and file operations.
How to read the score: The score is a causal estimate called IPS(τ̂), with 0 as the baseline. Higher scores indicate that using the model led to better outcomes. It combines five signals: user confirmation of success, praise or complaints, incorporation of correction requests, recovery from bash errors, and avoidance of tool hallucinations.
Caveats: The score scale differs from chat rankings, which use Bradley–Terry, so they cannot be compared directly. Models with fewer observations have wider confidence intervals.
Insights from this ranking
- The 95% confidence intervals for 1st-place Claude Fable 5.1 (Max) and 2nd-place Claude Opus 5.5 (High) overlap. This aggregation alone does not clearly establish their order.
- 4 other models have confidence intervals that overlap with the 1st-place model's. Interpret small ranking differences cautiously alongside vote counts.
- Anthropic has the most models among the top 10, with 6.
- The top-ranked model has 1,683,381 observations, and this leaderboard includes 50 models.
Overall agent performance LMArena
Published 2026-10-02 · Retrieved 2026-10-04
Scores use IPS(τ̂), with 0 as the baseline and higher scores indicating better performance. Lines show 95% confidence intervals. The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings.
Top ranking by company
- AnthropicClaude Fable 5.1 (Max)#1
- OOpenAIGPT 6 Astra (Max)#4
- GoogleGemini 4 Argon (High)#10
- Moonshot AIKimi K3 (Max)#14
- xxAIGrok 4.7 (xHigh)#16
- DeepSeekDeepseek V4.1 Flash (Max)#17
- MetaMuse Spark 1.3 (Max)#18
- TTencentHy4 preview#19
- ZZ.aiGLM 5.2 (Max)#20
View table (top 50)
| Rank | Model | Developer | Score | 95% confidence interval | Observations |
|---|---|---|---|---|---|
| 1 | #1Claude Fable 5.1 (Max) | Anthropic | 0.143 | 0.124–0.162 | 1,683,381 |
| 2 | #2Claude Opus 5.5 (High) | Anthropic | 0.138 | 0.116–0.16 | 754,031 |
| 3 | #3Claude Sonnet 5.5 (Max) | Anthropic | 0.125 | 0.094–0.156 | 483,211 |
| 4 | #4OGPT 6 Astra (Max) | OpenAI | 0.123 | 0.1–0.145 | 888,641 |
| 5 | #5OGPT 6.1 Sol (Max) | OpenAI | 0.112 | 0.085–0.14 | 379,372 |
| 6 | #6OGPT 6 Sol (Max) | OpenAI | 0.097 | 0.073–0.121 | 994,367 |
| 7 | #7Claude Opus 5 (High) | Anthropic | 0.087 | 0.073–0.101 | 3,531,760 |
| 8 | #8Claude Fable 5 (High) | Anthropic | 0.082 | 0.07–0.094 | 2,603,067 |
| 9 | #9Claude Opus 5 (Max) | Anthropic | 0.079 | 0.064–0.094 | 3,006,257 |
| 10 | #10Gemini 4 Argon (High) | 0.076 | 0.056–0.095 | 587,353 | |
| 11 | #11Claude Opus 4.8 (High) | Anthropic | 0.066 | 0.053–0.08 | 2,431,185 |
| 12 | #12OGPT 5.6 Sol (xHigh) | OpenAI | 0.065 | 0.052–0.077 | 4,348,715 |
| 13 | #13Claude Sonnet 5 (High) | Anthropic | 0.044 | 0.028–0.06 | 3,612,898 |
| 14 | #14Kimi K3 (Max) | Moonshot AI | 0.042 | 0.036–0.047 | 12,305,667 |
| 15 | #15OGPT 5.5 (xHigh) | OpenAI | 0.042 | 0.031–0.052 | 3,060,353 |
| 16 | #16xGrok 4.7 (xHigh) | xAI | 0.04 | 0.024–0.056 | 1,656,641 |
| 17 | #17Deepseek V4.1 Flash (Max) | DeepSeek | 0.04 | 0.036–0.045 | 16,932,092 |
| 18 | #18Muse Spark 1.3 (Max) | Meta | 0.04 | 0.034–0.046 | 6,139,426 |
| 19 | #19THy4 preview | Tencent | 0.04 | 0.032–0.047 | 5,108,255 |
| 20 | #20ZGLM 5.2 (Max) | Z.ai | 0.035 | 0.028–0.042 | 6,684,620 |
| 21 | #21Gemini 3.8 Flash (High) | 0.03 | 0.021–0.038 | 4,279,699 | |
| 22 | #22Qwen3.8 Max | Alibaba | 0.025 | 0.019–0.031 | 4,577,924 |
| 23 | #23ZGLM 5.3 (Max) | Z.ai | 0.024 | 0.017–0.03 | 9,155,315 |
| 24 | #24OGPT 6 Luna (Max) | OpenAI | 0.014 | 0.003–0.024 | 3,443,780 |
| 25 | #25xGrok 4.6 (xHigh) | xAI | 0.013 | 0.003–0.023 | 3,131,725 |
| 26 | #26xGrok 4.5 | xAI | 0.012 | 0.002–0.022 | 2,418,629 |
| 27 | #27DeepSeek V4 Pro (High) (0813) | DeepSeek | 0.012 | -0.005–0.028 | 797,284 |
| 28 | #28OGPT 5.4 (High) | OpenAI | 0.011 | 0.001–0.02 | 4,332,904 |
| 29 | #29OGPT 5.5 | OpenAI | 0.01 | 0.001–0.02 | 2,659,460 |
| 30 | #30SStep 5 Preview | StepFun | 0.005 | -0.01–0.02 | 1,125,598 |
| 31 | #31OGPT 5.6 Terra (xHigh) | OpenAI | 0.004 | -0.008–0.016 | 1,686,712 |
| 32 | #32ZGLM 5.3 Flash | Z.ai | -0.004 | -0.009–0 | 9,533,773 |
| 33 | #33MiMo V2.6 Flash | Xiaomi | -0.006 | -0.019–0.007 | 1,562,757 |
| 34 | #34Qwen3.8 Flash Next | Alibaba | -0.007 | -0.014–-0.001 | 11,249,221 |
| 35 | #35OGPT 5.6 Luna (xHigh) | OpenAI | -0.012 | -0.018–-0.005 | 4,028,029 |
| 36 | #36Gemini 3.7 Flash (High) | -0.015 | -0.022–-0.007 | 5,048,920 | |
| 37 | #37Qwen 3.8 27B | Alibaba | -0.017 | -0.023–-0.011 | 6,972,838 |
| 38 | #38Muse Spark 1.2 (xHigh) | Meta | -0.033 | -0.039–-0.026 | 2,884,900 |
| 39 | #39Muse Spark 1.1 | Meta | -0.049 | -0.053–-0.044 | 7,612,456 |
| 40 | #40Qwen3.7 Max | Alibaba | -0.051 | -0.061–-0.042 | 2,207,521 |
| 41 | #41THy3 | Tencent | -0.054 | -0.066–-0.042 | 1,171,568 |
| 42 | #42Minimax M3 | MiniMax | -0.069 | -0.078–-0.06 | 3,338,791 |
| 43 | #43Qwen3.7 Plus | Alibaba | -0.072 | -0.086–-0.057 | 1,144,063 |
| 44 | #44Mimo V2.5 Pro | Xiaomi | -0.075 | -0.084–-0.066 | 2,309,389 |
| 45 | #45Gemini 3.6 Flash (High) | -0.077 | -0.087–-0.066 | 1,523,777 | |
| 46 | #46Gemini 3.1 Pro Preview | -0.077 | -0.087–-0.067 | 3,446,977 | |
| 47 | #47TInkling Small | Thinky | -0.103 | -0.117–-0.089 | 503,653 |
| 48 | #48TInkling | Thinky | -0.109 | -0.119–-0.098 | 1,798,990 |
| 49 | #49Mistral Medium 3.5 | Mistral AI | -0.124 | -0.139–-0.109 | 437,621 |
| 50 | #50USolar Pro 4 | Upstage | -0.159 | -0.178–-0.141 | 427,705 |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings