Epoch AI · Fixing real-world code bugs (SWE-bench Verified)
We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.
How to read this leaderboard
What it measures: SWE-bench Verified evaluates the resolution of GitHub issues in real open-source Python repositories. The original dataset contains 500 issues from 12 repositories. Epoch's own evaluation uses 484 issues validated in its execution environment.
How to read the score: The percentage (%) of issues that pass both the specified solution verification tests and tests preserving existing functionality. Use it as a reference when comparing the ability to resolve issues in real repositories.
Caveats: Scores depend on tools and agent settings as well as the model. The evaluation focuses on Python repositories. Also check the number of issues evaluated and execution conditions.
Insights from this ranking
- The highest score is 83.5%, achieved by Claude Opus 4.7.
- The gap between 1st and 5th place is 4.8%p.
- Results for this test are public for 30 models.
Fixing real-world code bugs (SWE-bench Verified) Epoch AI
This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04
Bars show accuracy (%) and start at 0%.
View table (top 30)
| Rank | Model | Developer | Release date | Score |
|---|---|---|---|---|
| 1 | Claude Opus 4.7(max) | Anthropic | 2026-04-16 | 83.5% |
| 2 | OGPT-5.5(xhigh) | OpenAI | 2026-04-23 | 80.6% |
| 3 | Gemini 3.5 Flash | 2026-05-19 | 79.3% | |
| 4 | Claude Opus 4.6(no thinking) | Anthropic | 2026-02-05 | 78.7% |
| 5 | ZGLM-5.2 | Z.ai | 2026-06-16 | 78.7% |
| 6 | DeepSeek v4 Pro(max) | DeepSeek | 2026-04-24 | 77.6% |
| 7 | Qwen3.7 Max | Alibaba | 2026-05-19 | 77.3% |
| 8 | OGPT-5.4(high) | OpenAI | 2026-03-05 | 76.9% |
| 9 | Qwen 3.6 Max(Preview) | Alibaba | 2026-04-20 | 76.7% |
| 10 | Kimi K2.6 | Moonshot AI | 2026-04-20 | 76.7% |
| 11 | Claude Opus 4.5(no thinking) | Anthropic | 2025-11-24 | 76.7% |
| 12 | Gemini 3.1 Pro Preview | 2026-02-19 | 75.6% | |
| 13 | Gemini 3 Flash Preview | 2025-12-17 | 75.4% | |
| 14 | Claude Sonnet 4.6(no thinking) | Anthropic | 2026-02-17 | 75.2% |
| 15 | OGPT-5.3 Codex(high) | OpenAI | 2026-02-05 | 74.8% |
| 16 | ZGLM-5.1 | Z.ai | 2026-04-07 | 74.2% |
| 17 | Kimi K2.5 | Moonshot AI | 2026-01-27 | 73.8% |
| 18 | OGPT-5.2(high) | OpenAI | 2025-12-11 | 73.8% |
| 19 | OGPT-5(high) | OpenAI | 2025-08-07 | 73.6% |
| 20 | claude-opus-4-1-20250805 | Anthropic | 2025-08-05 | 73.3% |
| 21 | Gemini 3 Pro Preview | 2025-11-18 | 72.9% | |
| 22 | ZGLM-5 | Z.ai | 2026-02-11 | 72.1% |
| 23 | Claude Sonnet 4.5(no thinking) | Anthropic | 2025-09-29 | 71.3% |
| 24 | Claude Opus 4 | Anthropic | 2025-05-22 | 70.7% |
| 25 | OGPT-5.1(high) | OpenAI | 2025-11-13 | 68.0% |
| 26 | OGPT-5 mini(medium) | OpenAI | 2025-08-07 | 64.7% |
| 27 | Oo3(medium) | OpenAI | 2025-04-16 | 62.3% |
| 28 | claude-3-7-sonnet-20250219 | Anthropic | 2025-02-24 | 61.0% |
| 29 | Qwen 3.6 Plus(2026-04-02) | Alibaba | 2026-03-31 | 57.9% |
| 30 | gemini-2.5-pro | 2025-06-05 | 57.6% |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks