Valumigo

LMArena · Overall

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: This LMArena ranking aggregates user votes comparing responses from two models whose names are hidden. Like the LMArena website's default, it uses 'style and length adjustment (style control)' rankings that statistically adjust for the influence of response length and formatting, such as Markdown, on votes.

How to read the score: A relative score estimated from votes using the Bradley–Terry model and displayed on an Elo-like scale. The 95% confidence intervals indicate estimation uncertainty. Overlapping intervals alone do not establish equal capability.

Caveats: Because it evaluates which answers users prefer, results can be influenced by tone, length and formatting as well as accuracy. Also check the applied filters and adjustment settings.

Insights from this ranking

  • The 95% confidence intervals for 1st-place gemini-4-argon-high and 2nd-place claude-opus-4-6-high do not overlap. This aggregation provides relatively clear evidence that the 1st-place model leads.
  • Anthropic has the most models among the top 10, with 7.
  • The 1st-place model received 4,932 votes, and this leaderboard includes 413 models.

Overall LMArena

Published 2026-10-02 · Retrieved 2026-10-04

14601480150015201540
Googlegemini-4-argon-high
1525
Anthropicclaude-opus-4-6-high
1505
Anthropicclaude-fable-5-high
1504
Anthropicclaude-opus-5.5-high
1504
Anthropicclaude-opus-4-7-high
1501
Anthropicclaude-fable-5.1-max
1501
Googlegemini-3.8-flash-high
1495
Metamuse-spark-1.3-max
1494
Metamuse-spark-1.2 (xHigh)
1494
Metamuse-spark-1.1
1491
Anthropicclaude-opus-5-high
1490
Metamuse-spark
1489
Moonshot AIkimi-k3-max
1488
Googlegemini-3.7-flash-high
1488
Googlegemini-3.1-pro-preview
1487
Googlegemini-3-pro
1485
Ogpt-5.6-sol-xhigh
1484
Ogpt-6.1-sol-max
1483
Googlegemini-3.6-flash-high
1483
Alibabaqwen3.8-max
1482

Dots show scores; horizontal lines show 95% confidence intervals. Overlapping intervals suggest similar performance. The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings. Chat and vision rankings use 'style and length adjustment (style control)', the same default as the LMArena website.

Top ranking by company

  • GoogleGooglegemini-4-argon-high#1
  • AnthropicAnthropicclaude-opus-4-6-high#2
  • MetaMetamuse-spark-1.3-max#9
  • Moonshot AIMoonshot AIkimi-k3-max#16
  • OOpenAIgpt-5.6-sol-xhigh#20
  • AlibabaAlibabaqwen3.8-max#23
  • XiaomiXiaomimimo-v2.6-pro#26
  • ZZ.aiglm-5.3-max#27
  • xxAIgrok-4.20-beta1#36
View table (top 50)
RankModelDeveloperScore95% confidence intervalVotes
1#1Googlegemini-4-argon-highGoogle15251516–15344,932
2#2Anthropicclaude-opus-4-6-highAnthropic15051501–150877,636
3#3Anthropicclaude-fable-5-highAnthropic15041500–150938,387
4#4Anthropicclaude-opus-5.5-highAnthropic15041495–15134,552
5#5Anthropicclaude-opus-4-7-highAnthropic15011498–150564,946
6#6Anthropicclaude-fable-5.1-maxAnthropic15011495–150811,800
7#7Anthropicclaude-opus-4-6Anthropic14971494–150182,189
8#8Googlegemini-3.8-flash-highGoogle14951490–150026,298
9#9Metamuse-spark-1.3-maxMeta14941488–150012,343
10#10Anthropicclaude-opus-4-7Anthropic14941490–149866,058
11#11Metamuse-spark-1.2 (xHigh)Meta14941484–15034,356
12#12Metamuse-spark-1.1Meta14911487–149637,238
13#13Anthropicclaude-opus-5-highAnthropic14901486–149459,482
14#14Anthropicclaude-opus-5-maxAnthropic14891485–149428,937
15#15Metamuse-sparkMeta14891484–149514,120
16#16Moonshot AIkimi-k3-maxMoonshot AI14881484–149328,280
17#17Googlegemini-3.7-flash-highGoogle14881483–149321,539
18#18Googlegemini-3.1-pro-previewGoogle14871484–1490121,806
19#19Googlegemini-3-proGoogle14851482–148941,910
20#20Ogpt-5.6-sol-xhighOpenAI14841479–148836,656
21#21Ogpt-6.1-sol-maxOpenAI14831473–14943,071
22#22Googlegemini-3.6-flash-highGoogle14831479–148736,070
23#23Alibabaqwen3.8-maxAlibaba14821476–148723,353
24#24Anthropicclaude-opus-4-8-highAnthropic14821478–148563,974
25#25Ogpt-5.5-highOpenAI14811478–148569,202
26#26Xiaomimimo-v2.6-proXiaomi14801471–14894,056
27#27Zglm-5.3-maxZ.ai14781473–148417,857
28#28Googlegemini-3.5-flash-highGoogle14771473–148148,420
29#29Ogpt-6-astra-maxOpenAI14771470–14849,156
30#30Ogpt-5.5OpenAI14771473–148070,781
31#31Googlegemini-3.5-flash-mediumGoogle14761472–148047,187
32#32Zglm-5.2-maxZ.ai14761471–148045,399
33#33Ogpt-5.2-chat-latest-20260210OpenAI14761472–148035,937
34#34Ogpt-5.4-highOpenAI14751471–147964,642
35#35Alibabaqwen3.7-max-previewAlibaba14751465–14853,920
36#36xgrok-4.20-beta1xAI14751470–147927,827
37#37Anthropicclaude-opus-4-8Anthropic14751471–147865,072
38#38DeepSeekdeepseek-v4.1-flash-maxDeepSeek14741468–14818,728
39#39Anthropicclaude-opus-4-5-20251101-high-32kAnthropic14741470–147737,448
40#40Ogpt-5.5-instantOpenAI14731468–147827,098
41#41Zglm-5.3-flashZ.ai14731468–147822,971
42#42Googlegemini-3-flashGoogle14731468–147731,259
43#43Anthropicclaude-sonnet-4-6Anthropic14721469–147670,747
44#44xgrok-4.20-beta-0309-reasoningxAI14721468–147566,334
45#45Anthropicclaude-sonnet-5.5-xhighAnthropic14711461–14813,145
46#46xgrok-4.20-multi-agent-beta-0309xAI14711467–147464,788
47#47Anthropicclaude-opus-4-5-20251101Anthropic14701467–147372,772
48#48Baiduernie-5.1Baidu14681463–147239,396
49#49Xiaomimimo-v2.5-proXiaomi14681464–147170,871
50#50xgrok-4.5xAI14661461–147039,789

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings