Valumigo

Epoch AI · Fixing real-world code bugs (SWE-bench Verified)

We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: SWE-bench Verified evaluates the resolution of GitHub issues in real open-source Python repositories. The original dataset contains 500 issues from 12 repositories. Epoch's own evaluation uses 484 issues validated in its execution environment.

How to read the score: The percentage (%) of issues that pass both the specified solution verification tests and tests preserving existing functionality. Use it as a reference when comparing the ability to resolve issues in real repositories.

Caveats: Scores depend on tools and agent settings as well as the model. The evaluation focuses on Python repositories. Also check the number of issues evaluated and execution conditions.

Insights from this ranking

  • The highest score is 83.5%, achieved by Claude Opus 4.7.
  • The gap between 1st and 5th place is 4.8%p.
  • Results for this test are public for 30 models.

Fixing real-world code bugs (SWE-bench Verified) Epoch AI

This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04

AnthropicClaude Opus 4.7
83.5%
OGPT-5.5
80.6%
GoogleGemini 3.5 Flash
79.3%
AnthropicClaude Opus 4.6
78.7%
ZGLM-5.2
78.7%
DeepSeekDeepSeek v4 Pro
77.6%
AlibabaQwen3.7 Max
77.3%
OGPT-5.4
76.9%
AlibabaQwen 3.6 Max
76.7%
Moonshot AIKimi K2.6
76.7%
AnthropicClaude Opus 4.5
76.7%
GoogleGemini 3.1 Pro Preview
75.6%
GoogleGemini 3 Flash Preview
75.4%
AnthropicClaude Sonnet 4.6
75.2%
OGPT-5.3 Codex
74.8%
ZGLM-5.1
74.2%
Moonshot AIKimi K2.5
73.8%
OGPT-5.2
73.8%
OGPT-5
73.6%
Anthropicclaude-opus-4-1-20250805
73.3%

Bars show accuracy (%) and start at 0%.

View table (top 30)
RankModelDeveloperRelease dateScore
1AnthropicClaude Opus 4.7(max)Anthropic2026-04-1683.5%
2OGPT-5.5(xhigh)OpenAI2026-04-2380.6%
3GoogleGemini 3.5 FlashGoogle2026-05-1979.3%
4AnthropicClaude Opus 4.6(no thinking)Anthropic2026-02-0578.7%
5ZGLM-5.2Z.ai2026-06-1678.7%
6DeepSeekDeepSeek v4 Pro(max)DeepSeek2026-04-2477.6%
7AlibabaQwen3.7 MaxAlibaba2026-05-1977.3%
8OGPT-5.4(high)OpenAI2026-03-0576.9%
9AlibabaQwen 3.6 Max(Preview)Alibaba2026-04-2076.7%
10Moonshot AIKimi K2.6Moonshot AI2026-04-2076.7%
11AnthropicClaude Opus 4.5(no thinking)Anthropic2025-11-2476.7%
12GoogleGemini 3.1 Pro PreviewGoogle2026-02-1975.6%
13GoogleGemini 3 Flash PreviewGoogle2025-12-1775.4%
14AnthropicClaude Sonnet 4.6(no thinking)Anthropic2026-02-1775.2%
15OGPT-5.3 Codex(high)OpenAI2026-02-0574.8%
16ZGLM-5.1Z.ai2026-04-0774.2%
17Moonshot AIKimi K2.5Moonshot AI2026-01-2773.8%
18OGPT-5.2(high)OpenAI2025-12-1173.8%
19OGPT-5(high)OpenAI2025-08-0773.6%
20Anthropicclaude-opus-4-1-20250805Anthropic2025-08-0573.3%
21GoogleGemini 3 Pro PreviewGoogle2025-11-1872.9%
22ZGLM-5Z.ai2026-02-1172.1%
23AnthropicClaude Sonnet 4.5(no thinking)Anthropic2025-09-2971.3%
24AnthropicClaude Opus 4Anthropic2025-05-2270.7%
25OGPT-5.1(high)OpenAI2025-11-1368.0%
26OGPT-5 mini(medium)OpenAI2025-08-0764.7%
27Oo3(medium)OpenAI2025-04-1662.3%
28Anthropicclaude-3-7-sonnet-20250219Anthropic2025-02-2461.0%
29AlibabaQwen 3.6 Plus(2026-04-02)Alibaba2026-03-3157.9%
30Googlegemini-2.5-proGoogle2025-06-0557.6%

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks