Valumigo

LMArena · Creative writing

We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: A ranking based on votes classified as creative writing questions, such as fiction, poetry and copywriting, in LMArena text conversations.

How to read the score: A relative ranking reflecting which answers users preferred for creative writing questions. It does not separately score entertainment value or style objectively.

Caveats: To check whether a model suits your needs and style preferences, compare models directly using the same request.

Insights from this ranking

  • The 95% confidence intervals for 1st-place claude-opus-5.5-high and 2nd-place gemini-4-argon-high overlap. This aggregation alone does not clearly establish their order.
  • 3 other models have confidence intervals that overlap with the 1st-place model's. Interpret small ranking differences cautiously alongside vote counts.
  • Anthropic has the most models among the top 10, with 7.
  • The 1st-place model received 1,058 votes, and this leaderboard includes 411 models.

Creative writing LMArena

Published 2026-10-02 · Retrieved 2026-10-04

14401470150015301560
Anthropicclaude-opus-5.5-high
1532
Googlegemini-4-argon-high
1531
Anthropicclaude-opus-4-6-high
1506
Anthropicclaude-fable-5.1-max
1505
Anthropicclaude-fable-5-high
1495
Anthropicclaude-opus-5-max
1491
Googlegemini-3.7-flash-high
1491
Googlegemini-3.8-flash-high
1489
Anthropicclaude-opus-4-7-high
1486
Anthropicclaude-opus-5-high
1486
Anthropicclaude-opus-4-6
1484
Googlegemini-3-pro
1482
Googlegemini-3.1-pro-preview
1482
Anthropicclaude-opus-4-7
1481
Alibabaqwen3.8-max
1478
Googlegemini-3.5-flash-medium
1470
Googlegemini-3.5-flash-high
1468
Googlegemini-3.6-flash-high
1466
Xiaomimimo-v2.6-pro
1465
Alibabaqwen3.5-max-preview
1464

Dots show scores; horizontal lines show 95% confidence intervals. Overlapping intervals suggest similar performance.

View table (top 50)
RankModelDeveloperScore95% confidence intervalVotes
1Anthropicclaude-opus-5.5-highAnthropic15321512–15511,058
2Googlegemini-4-argon-highGoogle15311511–15501,097
3Anthropicclaude-opus-4-6-highAnthropic15061499–151214,085
4Anthropicclaude-fable-5.1-maxAnthropic15051493–15182,645
5Anthropicclaude-fable-5-highAnthropic14951488–15038,182
6Anthropicclaude-opus-5-maxAnthropic14911482–14996,526
7Googlegemini-3.7-flash-highGoogle14911481–15004,776
8Googlegemini-3.8-flash-highGoogle14891480–14985,697
9Anthropicclaude-opus-4-7-highAnthropic14861479–149311,812
10Anthropicclaude-opus-5-highAnthropic14861479–149313,263
11Anthropicclaude-opus-4-6Anthropic14841478–149114,240
12Googlegemini-3-proGoogle14821474–14906,510
13Googlegemini-3.1-pro-previewGoogle14821476–148722,309
14Anthropicclaude-opus-4-7Anthropic14811474–148812,043
15Alibabaqwen3.8-maxAlibaba14781468–14874,555
16Googlegemini-3.5-flash-mediumGoogle14701463–14789,445
17Googlegemini-3.5-flash-highGoogle14681460–14759,476
18Googlegemini-3.6-flash-highGoogle14661458–14747,580
19Xiaomimimo-v2.6-proXiaomi14651444–1487788
20Alibabaqwen3.5-max-previewAlibaba14641453–14753,311
21Zglm-5.2-maxZ.ai14641457–14718,979
22Metamuse-sparkMeta14601446–14742,040
23Metamuse-spark-1.3-maxMeta14581446–14712,631
24Zglm-5.3-maxZ.ai14571447–14674,034
25Googlegemini-3-flashGoogle14571447–14664,814
26Ogpt-5.5-highOpenAI14551448–146213,120
27Anthropicclaude-opus-4-8-highAnthropic14551448–146112,798
28Moonshot AIkimi-k3-maxMoonshot AI14541445–14635,694
29Googlegemini-2.5-proGoogle14531448–145917,583
30Zglm-5.1Z.ai14521446–145911,254
31Ogpt-5.5OpenAI14521446–145913,354
32Alibabaqwen3.7-max-previewAlibaba14511424–1477507
33Anthropicclaude-sonnet-5.5-xhighAnthropic14501427–1474658
34Metamuse-spark-1.2 (xHigh)Meta14501429–1471852
35Ogpt-5.6-sol-xhighOpenAI14481440–14567,537
36Anthropicclaude-opus-4-8Anthropic14461440–145312,836
37DeepSeekdeepseek-v4-proDeepSeek14461438–14539,721
38Anthropicclaude-opus-4-5-20251101-high-32kAnthropic14451437–14535,676
39Anthropicclaude-opus-4-5-20251101Anthropic14431437–145011,392
40Anthropicclaude-sonnet-4-5-20250929Anthropic14421436–144812,367
41DeepSeekdeepseek-v4-pro-high-previewDeepSeek14421434–14499,329
42Baiduernie-5.1Baidu14411432–14496,446
43xgrok-4.5xAI14411433–14488,291
44Zglm-5.3-flashZ.ai14411431–14504,870
45Ogpt-5.4-highOpenAI14391432–144611,109
46Zglm-5Z.ai14391430–14484,837
47Alibabaqwen3.7-plusAlibaba14391431–14477,805
48Metamuse-spark-1.1Meta14391431–14477,706
49Xiaomimimo-v2.5-proXiaomi14391432–144513,127
50xgrok-4.20-beta1xAI14381428–14484,158

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks