Epoch AI · Overall capability index (ECI)
We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.
How to read this leaderboard
What it measures: A composite capability index from Epoch AI that combines multiple benchmark scores on a common scale. It statistically estimates each benchmark's difficulty and how its scores change, enabling comparisons between models evaluated on different tests.
How to read the score: This is an estimate of relative capability, not an accuracy rate. Score differences cannot be directly converted into differences you would experience in actual use. Use it as a reference for overall standing across multiple tests.
Caveats: For capabilities needed for a specific use, also check the relevant benchmarks individually. Estimates and uncertainty may change as evaluation results or included benchmarks are added.
Insights from this ranking
- The highest score is 167.4, achieved by Claude Opus 5.5.
- The gap between 1st and 5th place is 4.5.
- Anthropic has the most models among the top 10, with 5.
- Results for this test are public for 60 models.
Overall capability index (ECI) Epoch AI
ECI combines results from multiple tests into an overall index created by Epoch AI. Higher scores indicate stronger overall capabilities. · Retrieved 2026-10-04
View table (top 40)
| Rank | Model | Developer | Release date | Score |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | Anthropic | 2026-09-22 | 167.4 |
| 2 | OGPT-6 Astra | OpenAI | 2026-09-03 | 166.5 |
| 3 | Claude Sonnet 5.5 | Anthropic | 2026-09-28 | 165.2 |
| 4 | Claude Fable 5.1 | Anthropic | 2026-09-01 | 164.8 |
| 5 | Claude Opus 5 | Anthropic | 2026-07-24 | 162.9 |
| 6 | OGPT-5.5 Pro | OpenAI | 2026-04-23 | 162.4 |
| 7 | Claude Fable 5 | Anthropic | 2026-06-09 | 162.2 |
| 8 | OGPT-5.6 Sol | OpenAI | 2026-07-09 | 161.8 |
| 9 | OGPT-5.6 Terra | OpenAI | 2026-07-09 | 159.8 |
| 10 | OGPT-5.5 | OpenAI | 2026-04-23 | 159.2 |
| 11 | OGPT-5.4 Pro | OpenAI | 2026-03-05 | 159.1 |
| 12 | Claude Opus 4.8 | Anthropic | 2026-05-28 | 158.3 |
| 13 | Kimi K3 | Moonshot AI | 2026-07-16 | 157.6 |
| 14 | Gemini 3.7 Flash | 2026-08-13 | 157.4 | |
| 15 | Gemini 3.8 Flash | 2026-09-02 | 156.9 | |
| 16 | OGPT-5.4 | OpenAI | 2026-03-05 | 156.9 |
| 17 | Muse Spark 1.3 | Meta | 2026-09-02 | 156.9 |
| 18 | OGPT-5.3 Codex | OpenAI | 2026-02-05 | 156.8 |
| 19 | xGrok 4.6 | xAI | 2026-08-12 | 156.6 |
| 20 | Qwen 3.8 Max | Alibaba | 2026-08-02 | 156.6 |
| 21 | OGPT-5.6 Luna | OpenAI | 2026-07-09 | 156.5 |
| 22 | Claude Opus 4.7 | Anthropic | 2026-04-16 | 156.4 |
| 23 | Claude Sonnet 5 | Anthropic | 2026-06-30 | 156.3 |
| 24 | ZGLM-5.3 | Z.ai | 2026-08-14 | 155.8 |
| 25 | OGPT-5.2 Pro | OpenAI | 2025-12-11 | 155.4 |
| 26 | DeepSeek V4 Pro 0813 | DeepSeek | 2026-08-13 | 155.4 |
| 27 | Claude Opus 4.6 | Anthropic | 2026-02-05 | 155.3 |
| 28 | Qwen3.8 Max (0902) | Alibaba | 2026-09-01 | 155.2 |
| 29 | DeepSeek V4.1 Flash | DeepSeek | 2026-09-09 | 155 |
| 30 | Muse Spark 1.2 | Meta | 2026-08-05 | 155 |
| 31 | Gemini 3.1 Pro | 2026-02-19 | 154.8 | |
| 32 | DeepSeek V4 Flash 0731 | DeepSeek | 2026-07-31 | 154.5 |
| 33 | Gemini 3.5 Flash | 2026-05-19 | 154.5 | |
| 34 | Gemini 3.6 Flash | 2026-07-21 | 154.3 | |
| 35 | Muse Spark 1.1 | Meta | 2026-07-09 | 154.3 |
| 36 | xGrok 4.5 | xAI | 2026-07-08 | 154 |
| 37 | Qwen3.7-Max | Alibaba | 2026-05-19 | 153.7 |
| 38 | OGPT-5.2 | OpenAI | 2025-12-11 | 153.5 |
| 39 | Gemini 3 Pro | 2025-11-18 | 153 | |
| 40 | Claude Sonnet 4.6 | Anthropic | 2026-02-17 | 152.3 |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks