Valumigo

Artificial Analysis · Agentic Index

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: An index created by Artificial Analysis by combining evaluations of multistep agent tasks involving tool use.

How to read the score: Higher scores indicate better results on tool use and multistep task evaluations.

Caveats: Results may vary with the evaluation environment and tool setup, so they should not be assumed to match performance in the agent tools you actually use.

Insights from this ranking

  • The highest score is 57.9, achieved by Claude Fable 5.1 (Max, Default Fallback).
  • The gap between 1st and 5th place is 4.9.
  • Anthropic has the most models among the top 10, with 3.

Agentic Index Artificial Analysis

Published 2026-10-05 · Retrieved 2026-10-05

AnthropicClaude Fable 5.1 (Max, Default Fallback)
57.9
AnthropicClaude Opus 5 (Max)
56.5
AlibabaQwen3.8 Max (0902)
56.0
ZGLM-5.3 (Max)
53.1
xGrok 4.6 (High)
53.0
OGPT-6 Astra (max)
51.0
ZGLM 5.3 Flash
50.9
AnthropicClaude Fable 5 (Max, Opus 4.8 Fallback)
50.7
OGPT-5.6 Sol (Max)
50.2
AlibabaQwen3.8 2.4T A95B
50.1
Moonshot AIKimi K3 (Max)
50.0
AlibabaQwen3.8 Max
49.9
AlibabaQwen3.8 27B (Xhigh)
45.8
AnthropicClaude Sonnet 5 (Max)
43.6
MetaMuse Spark 1.2 (Xhigh)
43.2
OGPT-5.6 Terra (Max)
43.2
OGPT-5.6 Luna (Max)
42.1
AnthropicClaude Opus 4.8 (Max)
41.9
DeepSeekDeepSeek V4 Pro 0813 (Max)
41.3
xGrok 4.5 (High)
41.2

The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings.

View table (top 60)
RankModelDeveloperScore
1#1AnthropicClaude Fable 5.1 (Max, Default Fallback)Anthropic57.9
2#2AnthropicClaude Opus 5 (Max)Anthropic56.5
3#3AlibabaQwen3.8 Max (0902)Alibaba56
4#4ZGLM-5.3 (Max)Z.ai53.1
5#5xGrok 4.6 (High)xAI53
6#6OGPT-6 Astra (max)OpenAI51
7#7ZGLM 5.3 FlashZ.ai50.9
8#8AnthropicClaude Fable 5 (Max, Opus 4.8 Fallback)Anthropic50.7
9#9OGPT-5.6 Sol (Max)OpenAI50.2
10#10AlibabaQwen3.8 2.4T A95BAlibaba50.1
11#11Moonshot AIKimi K3 (Max)Moonshot AI50
12#12AlibabaQwen3.8 MaxAlibaba49.9
13#13AlibabaQwen3.8 27B (Xhigh)Alibaba45.8
14#14AnthropicClaude Sonnet 5 (Max)Anthropic43.6
15#15MetaMuse Spark 1.2 (Xhigh)Meta43.2
16#16OGPT-5.6 Terra (Max)OpenAI43.2
17#17OGPT-5.6 Luna (Max)OpenAI42.1
18#18AnthropicClaude Opus 4.8 (Max)Anthropic41.9
19#19DeepSeekDeepSeek V4 Pro 0813 (Max)DeepSeek41.3
20#20xGrok 4.5 (High)xAI41.2
21#21DeepSeekDeepSeek V4 Flash 0731 (Max)DeepSeek41
22#22GoogleGemini 3.8 Flash (High)Google40.2
23#23AnthropicClaude Opus 4.7 (Max)Anthropic38.6
24#24ZGLM-5.2 (Max)Z.ai38.4
25#25OGPT-5.5 (Xhigh)OpenAI36.4
26#26GoogleGemini 3.7 Flash (High)Google35.2
27#27AnthropicClaude Sonnet 4.6 (Max)Anthropic31.8
28#28MiniMaxMiniMax-M3MiniMax29.5
29#29GoogleGemini 3.6 Flash (High)Google29
30#30iLing-3.0-flash-VLinclusionAI28.7
31#31iLing-3.0-flash-FininclusionAI27.9
32#32DeepSeekDeepSeek V4 Flash 0420 (High)DeepSeek26.3
33#33DeepSeekDeepSeek V4 Pro 0424 (Max)DeepSeek26.3
34#34GoogleGemini 3.5 Flash (High)Google26
35#35MetaMuse Spark 1.1 (Xhigh)Meta25.8
36#36THy3Tencent24.1
37#37ZGLM-5.1 (Reasoning)Z.ai23.9
38#38NNex-N2-Pro (Based on Qwen3.5-397B-A17B)Nex Agi23.7
39#39TInkling SmallThinkingmachines23.5
40#40TInkling (Xhigh)Thinkingmachines22.5
41#41AlibabaQwen3.7 MaxAlibaba22.5
42#42XiaomiMiMo-V2.5-Pro (Reasoning)Xiaomi21.3
43#43Moonshot AIKimi K2.7 CodeMoonshot AI21
44#44Moonshot AIKimi K2.6 (Reasoning)Moonshot AI20.5
45#45NVIDIANemotron 3 Ultra 550B A55B (Reasoning)NVIDIA20.1
46#46iLing 3.0 FlashinclusionAI19.3
47#47AlibabaQwen3.6 27B (Reasoning)Alibaba18.5
48#48OGPT-5.4 mini (Xhigh)OpenAI17.9
49#49AlibabaQwen3.7 PlusAlibaba17.5
50#50OGPT-5.4 nano (Xhigh)OpenAI16
51#51AnthropicClaude 4.5 Sonnet (Reasoning)Anthropic15.8
52#52XiaomiMiMo-V2.5Xiaomi15.8
53#53xGrok 4.3 (High)xAI15.5
54#54MiniMaxMiniMax-M2.7MiniMax15.3
55#55GoogleGemini 3.5 Flash-LiteGoogle14.3
56#56MLongCat 2.0Meituan14
57#57AlibabaQwen3.6 35B A3B (Reasoning)Alibaba13.1
58#58iRing-2.6-1TinclusionAI10.7
59#59AlibabaQwen3.5 397B A17B (Reasoning)Alibaba8.3
60#60GoogleGemini 3.1 Pro PreviewGoogle8.2

Source: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings). Licensed under CC BY 4.0.

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings). Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). Source: Design Arena (www.designarena.ai) via OpenRouter (openrouter.ai/rankings). Licensed under CC BY 4.0. Source: OpenRouter (openrouter.ai/rankings), as of 2026-10-05. Licensed under CC BY 4.0. https://openrouter.ai/rankings