Valumigo

LMArena · Japanese

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: A ranking calculated from votes on questions classified as Japanese in LMArena text conversations.

How to read the score: Shows which models voters preferred for questions in Japanese. The question's language does not indicate the voter's native language or identity.

Caveats: Vote counts and confidence interval widths vary by language and model. For new models, check vote counts, update dates and aggregation conditions together.

Insights from this ranking

  • The 95% confidence intervals for 1st-place claude-fable-5.1-max and 2nd-place claude-fable-5-high overlap. This aggregation alone does not clearly establish their order.
  • 25 other models have confidence intervals that overlap with the 1st-place model's. Interpret small ranking differences cautiously alongside vote counts.
  • Google has the most models among the top 10, with 4.
  • The 1st-place model received 200 votes, and this leaderboard includes 274 models.

Japanese LMArena

Published 2026-10-02 · Retrieved 2026-10-04

140014401480152015601600
Anthropicclaude-fable-5.1-max
1542
Anthropicclaude-fable-5-high
1524
Googlegemini-3-pro
1508
Anthropicclaude-opus-5-high
1506
Googlegemini-3.7-flash-high
1504
Ogpt-5.5-high
1503
Googlegemini-3.8-flash-high
1503
Moonshot AIkimi-k3-max
1502
Googlegemini-3.1-pro-preview
1499
Metamuse-spark-1.3-max
1498
Anthropicclaude-opus-4-6-high
1495
Ogpt-5.6-sol-xhigh
1495
Googlegemini-3-flash
1487
Alibabaqwen3.5-max-preview
1486
Anthropicclaude-opus-4-7-high
1485
Googlegemini-3.5-flash-high
1479
Ogpt-5.4-high
1479
xgrok-4.20-beta1
1478
Googlegemini-3.6-flash-high
1478
Alibabaqwen3.8-max
1473

Dots show scores; horizontal lines show 95% confidence intervals. Overlapping intervals suggest similar performance. The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings. Chat and vision rankings use 'style and length adjustment (style control)', the same default as the LMArena website.

Top ranking by company

  • AnthropicAnthropicclaude-fable-5.1-max#1
  • GoogleGooglegemini-3-pro#3
  • OOpenAIgpt-5.5-high#6
  • Moonshot AIMoonshot AIkimi-k3-max#8
  • MetaMetamuse-spark-1.3-max#10
  • AlibabaAlibabaqwen3.5-max-preview#15
  • xxAIgrok-4.20-beta1#20
  • DeepSeekDeepSeekdeepseek-v4-pro#31
  • ZZ.aiglm-5.2-max#33
View table (top 50)
RankModelDeveloperScore95% confidence intervalVotes
1#1Anthropicclaude-fable-5.1-maxAnthropic15421496–1588200
2#2Anthropicclaude-fable-5-highAnthropic15241496–1551538
3#3Googlegemini-3-proGoogle15081478–1538471
4#4Anthropicclaude-opus-5-highAnthropic15061483–1528958
5#5Googlegemini-3.7-flash-highGoogle15041471–1538340
6#6Ogpt-5.5-highOpenAI15031481–1525895
7#7Googlegemini-3.8-flash-highGoogle15031473–1533462
8#8Moonshot AIkimi-k3-maxMoonshot AI15021474–1530508
9#9Googlegemini-3.1-pro-previewGoogle14991481–15171,409
10#10Metamuse-spark-1.3-maxMeta14981452–1544188
11#11Anthropicclaude-opus-5-maxAnthropic14971467–1526482
12#12Anthropicclaude-opus-4-6-highAnthropic14951473–1518845
13#13Ogpt-5.6-sol-xhighOpenAI14951468–1523549
14#14Googlegemini-3-flashGoogle14871453–1521337
15#15Alibabaqwen3.5-max-previewAlibaba14861442–1531207
16#16Anthropicclaude-opus-4-7-highAnthropic14851461–1510705
17#17Anthropicclaude-opus-4-6Anthropic14801459–1502875
18#18Googlegemini-3.5-flash-highGoogle14791454–1504672
19#19Ogpt-5.4-highOpenAI14791454–1504671
20#20xgrok-4.20-beta1xAI14781437–1518225
21#21Googlegemini-3.6-flash-highGoogle14781452–1504598
22#22Ogpt-5.5OpenAI14771454–1501800
23#23Alibabaqwen3.8-maxAlibaba14731442–1505402
24#24Anthropicclaude-opus-4-7Anthropic14731450–1496797
25#25Metamuse-spark-1.1Meta14711445–1496633
26#26Ogpt-5.6-terra-xhighOpenAI14661439–1494559
27#27Googlegemini-3.5-flash-mediumGoogle14611436–1486668
28#28Ogpt-5.1-highOpenAI14601432–1488458
29#29Ogpt-5.5-instantOpenAI14601422–1498258
30#30Anthropicclaude-opus-4-8Anthropic14581436–1480874
31#31DeepSeekdeepseek-v4-proDeepSeek14571433–1481715
32#32Anthropicclaude-opus-4-8-highAnthropic14561434–1478881
33#33Zglm-5.2-maxZ.ai14551431–1480702
34#34Zglm-5.3-maxZ.ai14541420–1488327
35#35Ogpt-5.2-chat-latest-20260210OpenAI14491412–1485266
36#36Googlegemini-3.5-flash-liteGoogle14461419–1473565
37#37Googlegemini-2.5-proGoogle14451430–14602,037
38#38Ogpt-5.4OpenAI14451421–1469680
39#39Anthropicclaude-opus-4-5-20251101-high-32kAnthropic14441413–1476376
40#40Anthropicclaude-opus-4-5-20251101Anthropic14411417–1466633
41#41Moonshot AIkimi-k2.6Moonshot AI14411413–1469497
42#42Anthropicclaude-sonnet-5-highAnthropic14411416–1466662
43#43xgrok-4.5xAI14401412–1468522
44#44Zglm-4.7Z.ai14401385–1494123
45#45Anthropicclaude-sonnet-4-6Anthropic14401417–1462792
46#46Ogpt-4.5-preview-2025-02-27OpenAI14391401–1477271
47#47Ogpt-5.2-highOpenAI14371407–1467443
48#48Zglm-5.1Z.ai14361412–1459754
49#49Googlegemini-3-flash (thinking-minimal)Google14341412–1456847
50#50Zglm-5.3-flashZ.ai14331402–1464408

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings