Epoch AI · Reasoning through unfamiliar puzzles (ARC-AGI-2)
We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.
How to read this leaderboard
What it measures: ARC-AGI-2 was created by the ARC Prize Foundation. It requires inferring transformation rules from examples of input and output color grids and applying them to new inputs. The tasks are designed to be solvable by humans while remaining difficult for AI.
How to read the score: The percentage (%) of correctly predicted output grids. Evaluation uses pass@2, allowing two answer submissions per test input.
Caveats: Scores can vary with inference compute and solution methods. Check the evaluation set, attempt conditions and cost per task together.
Insights from this ranking
- The highest score is 95.0%, achieved by GPT-6 Astra.
- The gap between 1st and 5th place is 5.0%p.
- OpenAI has the most models among the top 10, with 4.
- Results for this test are public for 30 models.
Reasoning through unfamiliar puzzles (ARC-AGI-2) Epoch AI
This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04
Bars show accuracy (%) and start at 0%.
View table (top 30)
| Rank | Model | Developer | Release date | Score |
|---|---|---|---|---|
| 1 | OGPT-6 Astra(max) | OpenAI | 2026-09-03 | 95.0% |
| 2 | OGPT-5.6 Sol(max) | OpenAI | 2026-07-09 | 92.5% |
| 3 | Claude Opus 5.5(xhigh) | Anthropic | 2026-09-22 | 92.5% |
| 4 | Claude Opus 5(max) | Anthropic | 2026-07-24 | 90.4% |
| 5 | Claude Fable 5.1(max) | Anthropic | 2026-09-01 | 90.0% |
| 6 | OGPT-6 Sol(max) | OpenAI | 2026-09-22 | 89.6% |
| 7 | Claude Fable 5(max) | Anthropic | 2026-06-09 | 89.2% |
| 8 | OGPT-5.5(xhigh) | OpenAI | 2026-04-23 | 85.0% |
| 9 | Gemini 3.7 Flash(high) | 2026-08-13 | 84.6% | |
| 10 | Gemini 3 Deep Think | 2026-02-12 | 84.6% | |
| 11 | OGPT-5.5 Pro(high) | OpenAI | 2026-04-23 | 84.6% |
| 12 | OGPT-5.6 Terra(max) | OpenAI | 2026-07-09 | 83.9% |
| 13 | OGPT-5.4 Pro(xhigh) | OpenAI | 2026-03-05 | 83.3% |
| 14 | Gemini 3.1 Pro Preview | 2026-02-19 | 77.1% | |
| 15 | Claude Opus 4.7(max) | Anthropic | 2026-04-16 | 75.8% |
| 16 | OGPT-5.4(xhigh) | OpenAI | 2026-03-05 | 74.0% |
| 17 | Gemini 3.5 Flash | 2026-05-19 | 72.1% | |
| 18 | Claude Opus 4.8 | Anthropic | 2026-05-28 | 72.1% |
| 19 | Claude Opus 4.6(120k thinking) | Anthropic | 2026-02-05 | 69.2% |
| 20 | xGrok 4.6(xhigh) | xAI | 2026-08-12 | 67.1% |
| 21 | xgrok-4-20 | xAI | 2026-02-17 | 65.1% |
| 22 | DeepSeek V4 Flash 0731(max) | DeepSeek | 2026-07-31 | 61.4% |
| 23 | DeepSeek V4 Pro 0813(max) | DeepSeek | 2026-08-13 | 61.3% |
| 24 | Claude Sonnet 4.6(high) | Anthropic | 2026-02-17 | 60.4% |
| 25 | Kimi K3(max) | Moonshot AI | 2026-07-16 | 60.4% |
| 26 | Gemini 3.6 Flash | 2026-07-21 | 60.4% | |
| 27 | OGPT-5.6 Luna(max) | OpenAI | 2026-07-09 | 59.5% |
| 28 | OGPT-6 Luna(max) | OpenAI | 2026-09-22 | 59.3% |
| 29 | OGPT-5.2 Pro | OpenAI | 2025-12-11 | 54.2% |
| 30 | OGPT-5.2(xhigh) | OpenAI | 2025-12-11 | 52.9% |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks