Valumigo

LMArena · Math

We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: A ranking based on votes classified as math questions in LMArena text conversations.

How to read the score: A relative score reflecting which models users preferred for math answers. Models with fewer votes may have wider confidence intervals.

Caveats: This is a preference evaluation, not a test that directly scores answer accuracy. For mathematical accuracy, also check scored evaluations such as FrontierMath.

Insights from this ranking

  • The 95% confidence intervals for 1st-place gemini-4-argon-high and 2nd-place claude-opus-5-max overlap. This aggregation alone does not clearly establish their order.
  • 22 other models have confidence intervals that overlap with the 1st-place model's. Interpret small ranking differences cautiously alongside vote counts.
  • Anthropic has the most models among the top 10, with 7.
  • The 1st-place model received 255 votes, and this leaderboard includes 396 models.

Math LMArena

Published 2026-10-02 · Retrieved 2026-10-04

144014701500153015601590
Googlegemini-4-argon-high
1536
Anthropicclaude-opus-5-max
1531
Anthropicclaude-opus-5-high
1531
Googlegemini-3.8-flash-high
1523
Anthropicclaude-fable-5.1-max
1521
Anthropicclaude-fable-5-high
1518
Anthropicclaude-opus-4-6-high
1517
Anthropicclaude-opus-5.5-high
1516
Anthropicclaude-opus-4-6
1512
Googlegemini-3.7-flash-high
1507
Googlegemini-3.6-flash-high
1504
Googlegemini-3.5-flash-high
1504
Metamuse-spark-1.3-max
1502
Zglm-5.3-flash
1501
Alibabaqwen3.8-max
1500
Anthropicclaude-opus-4-7-high
1499
Xiaomimimo-v2.6-pro
1497
Moonshot AIkimi-k3-max
1494
Zglm-5.3-max
1491
Anthropicclaude-opus-4-7
1489

Dots show scores; horizontal lines show 95% confidence intervals. Overlapping intervals suggest similar performance.

View table (top 50)
RankModelDeveloperScore95% confidence intervalVotes
1Googlegemini-4-argon-highGoogle15361500–1572255
2Anthropicclaude-opus-5-maxAnthropic15311514–15471,337
3Anthropicclaude-opus-5-highAnthropic15311519–15432,756
4Googlegemini-3.8-flash-highGoogle15231506–15401,274
5Anthropicclaude-fable-5.1-maxAnthropic15211496–1545544
6Anthropicclaude-fable-5-highAnthropic15181504–15321,889
7Anthropicclaude-opus-4-6-highAnthropic15171507–15273,978
8Anthropicclaude-opus-5.5-highAnthropic15161480–1552235
9Anthropicclaude-opus-4-6Anthropic15121503–15224,454
10Googlegemini-3.7-flash-highGoogle15071489–15251,097
11Googlegemini-3.6-flash-highGoogle15041490–15181,736
12Googlegemini-3.5-flash-highGoogle15041491–15162,361
13Metamuse-spark-1.3-maxMeta15021478–1526586
14Zglm-5.3-flashZ.ai15011483–15191,131
15Alibabaqwen3.8-maxAlibaba15001482–15171,151
16Anthropicclaude-opus-4-7-highAnthropic14991488–15093,260
17Xiaomimimo-v2.6-proXiaomi14971457–1537213
18Moonshot AIkimi-k3-maxMoonshot AI14941477–15111,218
19Zglm-5.3-maxZ.ai14911470–1512739
20Anthropicclaude-opus-4-7Anthropic14891479–15003,371
21Alibabaqwen3.7-max-previewAlibaba14881450–1526231
22Ogpt-5.4-highOpenAI14881477–14983,379
23Googlegemini-3.5-flash-mediumGoogle14871474–14992,196
24Anthropicclaude-opus-4-8-highAnthropic14861475–14982,941
25DeepSeekdeepseek-v4.1-flash-maxDeepSeek14861459–1513433
26Ogpt-5.5OpenAI14861476–14973,499
27Googlegemini-3.1-pro-previewGoogle14851476–14936,301
28Zglm-5.2-maxZ.ai14841471–14972,035
29Xiaomimimo-v2.5-proXiaomi14821471–14923,356
30Baiduernie-5.1Baidu14811467–14941,980
31Metamuse-spark-1.1Meta14801466–14951,661
32Xiaomimimo-v2.6-flashXiaomi14791448–1510330
33TinklingThinky14781464–14931,581
34Ogpt-5.5-highOpenAI14781468–14893,472
35Googlegemini-3-proGoogle14751464–14862,761
36Moonshot AIkimi-k2.6Moonshot AI14751462–14882,075
37Alibabaqwen3.5-max-previewAlibaba14741458–14901,421
38Thy3Tencent14731448–1498550
39Zglm-5.1Z.ai14731461–14842,860
40Googlegemini-3-flashGoogle14721460–14852,060
41Ogpt-5.6-sol-xhighOpenAI14721457–14861,732
42Anthropicclaude-opus-4-8Anthropic14701459–14813,072
43Moonshot AIkimi-k2.5-thinkingMoonshot AI14701461–14804,158
44Ogpt-6-astra-maxOpenAI14691442–1496458
45Googlegemma-4-26b-a4bGoogle14691441–1496389
46Anthropicclaude-sonnet-5-highAnthropic14671454–14802,061
47Alibabaqwen3.7-plusAlibaba14661453–14801,915
48Alibabaqwen3.6-max-previewAlibaba14651436–1494380
49Googlegemma-4-31bGoogle14641438–1491432
50Anthropicclaude-sonnet-4-6Anthropic14631453–14733,778

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks