Valumigo

Design Arena · Game development

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: Rankings based on people's comparison votes on simple games created from the same prompt in Design Arena.

How to read the score: These are relative scores using the Elo method. Higher scores mean people preferred the results more often. Win rates are also shown.

Caveats: The evaluations focus on short, standalone games, so results may differ from the ability to develop large projects.

Insights from this ranking

  • The highest score is 1479, achieved by Claude Opus 5.5.
  • Anthropic has the most models among the top 10, with 4.

Game development Design Arena

Published 2026-10-05 · Retrieved 2026-10-05

125013001350140014501500
AnthropicClaude Opus 5.5
1479
AnthropicClaude Fable 5.1
1394
Moonshot AIKimi K3
1379
MetaMuse Spark 1.3 Max
1358
AnthropicClaude Opus 5
1350
AnthropicClaude Fable 5
1349
MetaMuse Spark 1.3 (xhigh)
1343
GoogleGemini 3.7 Flash
1335
ZGLM-5.3
1334
AlibabaQwen3.8 Max
1329
XiaomiMiMo-V2.6-Pro
1323
GoogleGemini 3.8 Flash
1320
MetaMuse Spark 1.2
1314
OGPT-5.5
1308
xGrok 4.6
1307
xGrok 4.7
1304
AnthropicClaude Sonnet 5
1296
ZGLM-5.3-Flash
1294
AnthropicClaude Opus 4.6 (Thinking)
1290
AnthropicClaude Opus 4.6
1285

The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings.

View table (top 60)
RankModelDeveloperScoreWin rate
1#1AnthropicClaude Opus 5.5Anthropic147973.6%
2#2AnthropicClaude Fable 5.1Anthropic139465.6%
3#3Moonshot AIKimi K3Moonshot AI137961.8%
4#4MetaMuse Spark 1.3 MaxMeta135856.2%
5#5AnthropicClaude Opus 5Anthropic135059.1%
6#6AnthropicClaude Fable 5Anthropic134960.8%
7#7MetaMuse Spark 1.3 (xhigh)Meta134354.6%
8#8GoogleGemini 3.7 FlashGoogle133554.2%
9#9ZGLM-5.3Z.ai133454.6%
10#10AlibabaQwen3.8 MaxAlibaba132953.3%
11#11XiaomiMiMo-V2.6-ProXiaomi132353%
12#12GoogleGemini 3.8 FlashGoogle132051.3%
13#13MetaMuse Spark 1.2Meta131450.4%
14#14OGPT-5.5OpenAI130858.1%
15#15xGrok 4.6xAI130754.2%
16#16xGrok 4.7xAI130451.5%
17#17AnthropicClaude Sonnet 5Anthropic129653%
18#18ZGLM-5.3-FlashZ.ai129447.2%
19#19AnthropicClaude Opus 4.6 (Thinking)Anthropic129061%
20#20AnthropicClaude Opus 4.6Anthropic128560.5%
21#21MetaMuse Spark 1.1Meta128548.1%
22#22ZGLM 5.2Z.ai128149.5%
23#23AlibabaQwen3.7 MaxAlibaba128053.6%
24#24AnthropicClaude Opus 4.8Anthropic127553.2%
25#25ZGLM 5.1Z.ai127555.6%
26#26GoogleGemini 3.6 FlashGoogle127551.1%
27#27xGrok 4.5xAI127549.3%
28#28GoogleGemini 3.5 FlashGoogle127453%
29#29AlibabaQwen3.7 PlusAlibaba127349.7%
30#30AnthropicClaude Sonnet 4.6Anthropic127258.8%
31#31XiaomiMiMo-V2.5-ProXiaomi127253.4%
32#32Moonshot AIKimi K2.6Moonshot AI126354.9%
33#33ZGLM 5 TurboZ.ai126053.2%
34#34XiaomiMiMo-V2.5Xiaomi125752.6%
35#35OGPT-5.4 (Medium)OpenAI125157.6%
36#36ZGLM 5Z.ai124757.4%
37#37AnthropicClaude Opus 4.5Anthropic124259.3%
38#38DeepSeekDeepSeek-V4-ProDeepSeek124252.3%
39#39NNex N2 ProNex Agi123849%
40#40Moonshot AIKimi K2.7 CodeMoonshot AI123347.5%
41#41ZGLM 5V TurboZ.ai123150.7%
42#42MiniMaxMiniMax M3MiniMax122546%
43#43AlibabaQwen3.6 PlusAlibaba122449.1%
44#44MiniMaxMiniMax M2.7MiniMax122350%
45#45DeepSeekDeepSeek-V4-Flash-0731DeepSeek122144.6%
46#46Moonshot AIKimi K2.5 (Thinking)Moonshot AI121953.4%
47#47GoogleGemini 3.1 Pro PreviewGoogle121552.1%
48#48DeepSeekDeepSeek-V4-FlashDeepSeek121050.2%
49#49GoogleGemini 3 Pro PreviewGoogle120960.9%
50#50OGPT-5.2 (Low)OpenAI120856%
51#51OGPT-5.4 (None)OpenAI120753.6%
52#52xGrok 4.20 Beta (Reasoning)xAI120550.1%
53#53OGPT-5.4 (Low)OpenAI120452.8%
54#54AnthropicClaude 3.7 SonnetAnthropic120354.3%
55#55ZGLM 4.7Z.ai120155.2%
56#56OGPT-5.2 (High)OpenAI119956.7%
57#57OGPT-5 (High)OpenAI119859.3%
58#58XiaomiMiMo-V2-FlashXiaomi119748.6%
59#59OGPT-5.2 (Medium)OpenAI119652.8%
60#60OGPT-5.1 (High)OpenAI119155.9%

Source: Design Arena (www.designarena.ai) via OpenRouter (openrouter.ai/rankings). Licensed under CC BY 4.0.

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings). Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). Source: Design Arena (www.designarena.ai) via OpenRouter (openrouter.ai/rankings). Licensed under CC BY 4.0. Source: OpenRouter (openrouter.ai/rankings), as of 2026-10-05. Licensed under CC BY 4.0. https://openrouter.ai/rankings