Valumigo

Epoch AI · Autonomous computer tasks (Terminal-Bench)

We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: Terminal-Bench 2.0 evaluates completing multistep tasks in a terminal environment. It covers command-line work such as using programs, installation, configuration and debugging.

How to read the score: The percentage (%) of tasks that pass the scoring tests, sometimes reported as an average across repeated runs. It evaluates combinations of models and tools such as Claude Code and Codex CLI.

Caveats: Scores vary with agent tools and settings. Epoch's representative value for each model is its highest-scoring model-agent combination, so check the execution tools as well.

Insights from this ranking

  • The highest score is 84.7%, achieved by GPT-5.5.
  • The gap between 1st and 5th place is 4.9%p.
  • OpenAI has the most models among the top 10, with 5.
  • Results for this test are public for 30 models.

Autonomous computer tasks (Terminal-Bench) Epoch AI

This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04

OGPT-5.5
84.7%
OGPT-5.4
81.8%
GoogleGemini 3.1 Pro Preview
80.2%
AnthropicClaude Opus 4.7
80.2%
AnthropicClaude Opus 4.6
79.8%
OGPT-5.3 Codex
78.4%
GoogleGemini 3 Pro Preview
69.4%
OGPT-5.2 Codex
66.5%
OGPT-5.2
64.9%
GoogleGemini 3 Flash Preview
64.3%
AnthropicClaude Opus 4.5
63.1%
OGPT-5.1 Codex Mini
61.6%
?
61.2%
OGPT-5.1-Codex-Max
60.4%
OGPT-5.1 Codex
57.8%
xgrok-4-20
57.3%
AnthropicClaude Sonnet 4.6
53.4%
ZGLM-5
52.4%
OGPT-5
49.6%
OGPT-5.1
47.6%

Bars show accuracy (%) and start at 0%.

View table (top 30)
RankModelDeveloperRelease dateScore
1OGPT-5.5(unknown thinking)OpenAI2026-04-2384.7%
2OGPT-5.4(unknown thinking)OpenAI2026-03-0581.8%
3GoogleGemini 3.1 Pro PreviewGoogle2026-02-1980.2%
4AnthropicClaude Opus 4.7(unknown)Anthropic2026-04-1680.2%
5AnthropicClaude Opus 4.6(unknown thinking)Anthropic2026-02-0579.8%
6OGPT-5.3 CodexOpenAI2026-02-0578.4%
7GoogleGemini 3 Pro PreviewGoogle2025-11-1869.4%
8OGPT-5.2 CodexOpenAI2025-12-1866.5%
9OGPT-5.2(unknown thinking)OpenAI2025-12-1164.9%
10GoogleGemini 3 Flash PreviewGoogle2025-12-1764.3%
11AnthropicClaude Opus 4.5(unknown thinking)Anthropic2025-11-2463.1%
12OGPT-5.1 Codex MiniOpenAI2025-11-1261.6%
13??61.2%
14OGPT-5.1-Codex-MaxOpenAI2025-11-1960.4%
15OGPT-5.1 CodexOpenAI2025-11-1257.8%
16xgrok-4-20xAI2026-02-1757.3%
17AnthropicClaude Sonnet 4.6(unknown thinking)Anthropic2026-02-1753.4%
18ZGLM-5Z.ai2026-02-1152.4%
19OGPT-5(unknown thinking)OpenAI2025-08-0749.6%
20OGPT-5.1(medium)OpenAI2025-11-1347.6%
21AnthropicClaude Sonnet 4.5(unknown thinking)Anthropic2025-09-2946.5%
22MiniMaxMiniMax-M2.7MiniMax2026-03-1845.1%
23OGPT-5-codexOpenAI2025-09-1544.3%
24Moonshot AIKimi K2.5Moonshot AI2026-01-2743.2%
25MiniMaxMiniMax-M2.5MiniMax2026-02-1242.7%
26DeepSeekDeepSeek-V3.2(Thinking; Novita)DeepSeek2025-12-0139.6%
27Anthropicclaude-opus-4-1-20250805Anthropic2025-08-0538.0%
28AnthropicClaude Opus 4.1(unknown thinking)Anthropic2025-08-0538.0%
29MiniMaxMiniMax-M2.1MiniMax2025-12-2336.6%
30Moonshot AIKimi K2 ThinkingMoonshot AI2025-11-0635.7%

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks