Design Arena · Game development
We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.
How to read this leaderboard
What it measures: Rankings based on people's comparison votes on simple games created from the same prompt in Design Arena.
How to read the score: These are relative scores using the Elo method. Higher scores mean people preferred the results more often. Win rates are also shown.
Caveats: The evaluations focus on short, standalone games, so results may differ from the ability to develop large projects.
Insights from this ranking
- The highest score is 1479, achieved by Claude Opus 5.5.
- Anthropic has the most models among the top 10, with 4.
Game development Design Arena
Published 2026-10-05 · Retrieved 2026-10-05
The chart shows only the highest-scoring reasoning setting, such as high or max, for each model. The table shows rankings for all settings.
View table (top 60)
| Rank | Model | Developer | Score | Win rate |
|---|---|---|---|---|
| 1 | #1Claude Opus 5.5 | Anthropic | 1479 | 73.6% |
| 2 | #2Claude Fable 5.1 | Anthropic | 1394 | 65.6% |
| 3 | #3Kimi K3 | Moonshot AI | 1379 | 61.8% |
| 4 | #4Muse Spark 1.3 Max | Meta | 1358 | 56.2% |
| 5 | #5Claude Opus 5 | Anthropic | 1350 | 59.1% |
| 6 | #6Claude Fable 5 | Anthropic | 1349 | 60.8% |
| 7 | #7Muse Spark 1.3 (xhigh) | Meta | 1343 | 54.6% |
| 8 | #8Gemini 3.7 Flash | 1335 | 54.2% | |
| 9 | #9ZGLM-5.3 | Z.ai | 1334 | 54.6% |
| 10 | #10Qwen3.8 Max | Alibaba | 1329 | 53.3% |
| 11 | #11MiMo-V2.6-Pro | Xiaomi | 1323 | 53% |
| 12 | #12Gemini 3.8 Flash | 1320 | 51.3% | |
| 13 | #13Muse Spark 1.2 | Meta | 1314 | 50.4% |
| 14 | #14OGPT-5.5 | OpenAI | 1308 | 58.1% |
| 15 | #15xGrok 4.6 | xAI | 1307 | 54.2% |
| 16 | #16xGrok 4.7 | xAI | 1304 | 51.5% |
| 17 | #17Claude Sonnet 5 | Anthropic | 1296 | 53% |
| 18 | #18ZGLM-5.3-Flash | Z.ai | 1294 | 47.2% |
| 19 | #19Claude Opus 4.6 (Thinking) | Anthropic | 1290 | 61% |
| 20 | #20Claude Opus 4.6 | Anthropic | 1285 | 60.5% |
| 21 | #21Muse Spark 1.1 | Meta | 1285 | 48.1% |
| 22 | #22ZGLM 5.2 | Z.ai | 1281 | 49.5% |
| 23 | #23Qwen3.7 Max | Alibaba | 1280 | 53.6% |
| 24 | #24Claude Opus 4.8 | Anthropic | 1275 | 53.2% |
| 25 | #25ZGLM 5.1 | Z.ai | 1275 | 55.6% |
| 26 | #26Gemini 3.6 Flash | 1275 | 51.1% | |
| 27 | #27xGrok 4.5 | xAI | 1275 | 49.3% |
| 28 | #28Gemini 3.5 Flash | 1274 | 53% | |
| 29 | #29Qwen3.7 Plus | Alibaba | 1273 | 49.7% |
| 30 | #30Claude Sonnet 4.6 | Anthropic | 1272 | 58.8% |
| 31 | #31MiMo-V2.5-Pro | Xiaomi | 1272 | 53.4% |
| 32 | #32Kimi K2.6 | Moonshot AI | 1263 | 54.9% |
| 33 | #33ZGLM 5 Turbo | Z.ai | 1260 | 53.2% |
| 34 | #34MiMo-V2.5 | Xiaomi | 1257 | 52.6% |
| 35 | #35OGPT-5.4 (Medium) | OpenAI | 1251 | 57.6% |
| 36 | #36ZGLM 5 | Z.ai | 1247 | 57.4% |
| 37 | #37Claude Opus 4.5 | Anthropic | 1242 | 59.3% |
| 38 | #38DeepSeek-V4-Pro | DeepSeek | 1242 | 52.3% |
| 39 | #39NNex N2 Pro | Nex Agi | 1238 | 49% |
| 40 | #40Kimi K2.7 Code | Moonshot AI | 1233 | 47.5% |
| 41 | #41ZGLM 5V Turbo | Z.ai | 1231 | 50.7% |
| 42 | #42MiniMax M3 | MiniMax | 1225 | 46% |
| 43 | #43Qwen3.6 Plus | Alibaba | 1224 | 49.1% |
| 44 | #44MiniMax M2.7 | MiniMax | 1223 | 50% |
| 45 | #45DeepSeek-V4-Flash-0731 | DeepSeek | 1221 | 44.6% |
| 46 | #46Kimi K2.5 (Thinking) | Moonshot AI | 1219 | 53.4% |
| 47 | #47Gemini 3.1 Pro Preview | 1215 | 52.1% | |
| 48 | #48DeepSeek-V4-Flash | DeepSeek | 1210 | 50.2% |
| 49 | #49Gemini 3 Pro Preview | 1209 | 60.9% | |
| 50 | #50OGPT-5.2 (Low) | OpenAI | 1208 | 56% |
| 51 | #51OGPT-5.4 (None) | OpenAI | 1207 | 53.6% |
| 52 | #52xGrok 4.20 Beta (Reasoning) | xAI | 1205 | 50.1% |
| 53 | #53OGPT-5.4 (Low) | OpenAI | 1204 | 52.8% |
| 54 | #54Claude 3.7 Sonnet | Anthropic | 1203 | 54.3% |
| 55 | #55ZGLM 4.7 | Z.ai | 1201 | 55.2% |
| 56 | #56OGPT-5.2 (High) | OpenAI | 1199 | 56.7% |
| 57 | #57OGPT-5 (High) | OpenAI | 1198 | 59.3% |
| 58 | #58MiMo-V2-Flash | Xiaomi | 1197 | 48.6% |
| 59 | #59OGPT-5.2 (Medium) | OpenAI | 1196 | 52.8% |
| 60 | #60OGPT-5.1 (High) | OpenAI | 1191 | 55.9% |
Source: Design Arena (www.designarena.ai) via OpenRouter (openrouter.ai/rankings). Licensed under CC BY 4.0.
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: Artificial Analysis (artificialanalysis.ai) via OpenRouter (openrouter.ai/rankings). Source: OpenRouter evals (openrouter.ai) via OpenRouter (openrouter.ai/rankings). Source: Design Arena (www.designarena.ai) via OpenRouter (openrouter.ai/rankings). Licensed under CC BY 4.0. Source: OpenRouter (openrouter.ai/rankings), as of 2026-10-05. Licensed under CC BY 4.0. https://openrouter.ai/rankings