Valumigo

LMArena · Document understanding

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: This LMArena ranking is based on votes comparing two models' responses to questions about uploaded documents, such as PDFs.

How to read the score: These are relative scores estimated using Bradley–Terry. Interpret ranking differences cautiously when 95% confidence intervals are wide.

Caveats: Updates may be less frequent than for other leaderboards, so check the publication date. Results may vary by document type and language.

Insights from this ranking

  • The 95% confidence intervals for 1st-place claude-opus-5-high and 2nd-place claude-fable-5.1-max overlap. This aggregation alone does not clearly establish their order.
  • 3 other models have confidence intervals that overlap with the 1st-place model's. Interpret small ranking differences cautiously alongside vote counts.
  • Anthropic has the most models among the top 10, with 7.
  • The 1st-place model received 8,794 votes, and this leaderboard includes 44 models.

Document understanding LMArena

Published 2026-09-13 · Retrieved 2026-10-04

1440147015001530
Anthropicclaude-opus-5-high
1516
Anthropicclaude-fable-5.1-max
1513
Anthropicclaude-opus-4-6-high
1507
Anthropicclaude-fable-5
1496
Anthropicclaude-opus-4-7
1495
Ogpt-5.5
1486
Ogpt-5.6-sol-xhigh
1483
Anthropicclaude-sonnet-4-6
1482
Anthropicclaude-opus-4-8-high
1475
Ogpt-5.6-terra-xhigh
1472
Ogpt-5.4
1471
Metamuse-spark-1.3-max
1471
Ogpt-6-astra-max
1468
Anthropicclaude-sonnet-5-high
1466
Metamuse-spark-1.1
1465
Googlegemini-3.5-flash-medium
1463
Anthropicclaude-opus-4-5-20251101
1462
Ogpt-5.6-luna-xhigh
1457
Googlegemini-3.6-flash-high
1456
xgrok-4.6-high
1452

Dots show scores; horizontal lines show 95% confidence intervals. Overlapping intervals suggest similar performance. The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings.

Top ranking by company

  • AnthropicAnthropicclaude-opus-5-high#1
  • OOpenAIgpt-5.5#8
  • MetaMetamuse-spark-1.3-max#15
  • GoogleGooglegemini-3.5-flash-medium#20
  • xxAIgrok-4.6-high#25
  • Moonshot AIMoonshot AIkimi-k2.6#27
  • AlibabaAlibabaqwen3.7-plus#30
  • MiniMaxMiniMaxminimax-m3#32
  • ZZ.aiglm-5v-turbo#39
View table (top 50)
RankModelDeveloperScore95% confidence intervalVotes
1#1Anthropicclaude-opus-5-highAnthropic15161509–15238,794
2#2Anthropicclaude-fable-5.1-maxAnthropic15131497–15281,403
3#3Anthropicclaude-opus-4-6-highAnthropic15071501–151427,929
4#4Anthropicclaude-opus-4-6Anthropic15071501–151341,260
5#5Anthropicclaude-fable-5Anthropic14961488–15048,497
6#6Anthropicclaude-opus-4-7Anthropic14951489–150122,192
7#7Anthropicclaude-opus-4-7-highAnthropic14951488–150121,957
8#8Ogpt-5.5OpenAI14861479–149221,279
9#9Ogpt-5.5-highOpenAI14841478–149120,944
10#10Ogpt-5.6-sol-xhighOpenAI14831474–14924,818
11#11Anthropicclaude-sonnet-4-6Anthropic14821476–148857,813
12#12Anthropicclaude-opus-4-8-highAnthropic14751468–148211,551
13#13Ogpt-5.6-terra-xhighOpenAI14721463–14805,787
14#14Ogpt-5.4OpenAI14711465–147733,331
15#15Metamuse-spark-1.3-maxMeta14711452–14891,006
16#16Ogpt-6-astra-maxOpenAI14681453–14821,621
17#17Anthropicclaude-sonnet-5-highAnthropic14661458–14747,411
18#18Metamuse-spark-1.1Meta14651456–14744,885
19#19Anthropicclaude-opus-4-8Anthropic14641457–147112,128
20#20Googlegemini-3.5-flash-mediumGoogle14631455–14716,500
21#21Anthropicclaude-opus-4-5-20251101Anthropic14621452–14737,981
22#22Googlegemini-3.5-flash-highGoogle14611452–14714,061
23#23Ogpt-5.6-luna-xhighOpenAI14571449–14655,766
24#24Googlegemini-3.6-flash-highGoogle14561445–14672,975
25#25xgrok-4.6-highxAI14521440–14642,258
26#26xgrok-4.5xAI14521443–14614,965
27#27Moonshot AIkimi-k2.6Moonshot AI14511443–145811,291
28#28Anthropicclaude-sonnet-4-5-20250929Anthropic14501444–145631,455
29#29Metamuse-sparkMeta14441426–14621,084
30#30Alibabaqwen3.7-plusAlibaba14441435–14534,591
31#31Googlegemini-3.1-pro-previewGoogle14441439–144949,457
32#32MiniMaxminimax-m3MiniMax14351427–14436,314
33#33Googlegemini-3-proGoogle14341425–144310,769
34#34Moonshot AIkimi-k2.5-thinkingMoonshot AI14301423–143721,569
35#35Googlegemma-4-31bGoogle14251417–143212,326
36#36Googlegemini-2.5-proGoogle14211415–142725,110
37#37Anthropicclaude-haiku-4-5-20251001Anthropic14201414–142634,677
38#38xgrok-4.20-beta-0309-reasoningxAI14161409–142320,948
39#39Zglm-5v-turboZ.ai14161407–14246,995
40#40Googlegemini-3-flashGoogle14131404–14227,158
41#41Ogpt-5.2-highOpenAI14051395–14147,108
42#42Ogpt-5.1OpenAI14031394–14138,244
43#43Ogpt-5.5-instantOpenAI14031395–14118,497
44#44Ogpt-5.2OpenAI14011395–140728,212

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings