Epoch AI · Autonomous computer tasks (Terminal-Bench)
We regularly fetch and display public benchmark data: LMArena user-voted rankings and Epoch AI test scores. Each leaderboard shows its publication and retrieval dates. These scores are not produced by this site.
How to read this leaderboard
What it measures: Terminal-Bench 2.0 evaluates completing multistep tasks in a terminal environment. It covers command-line work such as using programs, installation, configuration and debugging.
How to read the score: The percentage (%) of tasks that pass the scoring tests, sometimes reported as an average across repeated runs. It evaluates combinations of models and tools such as Claude Code and Codex CLI.
Caveats: Scores vary with agent tools and settings. Epoch's representative value for each model is its highest-scoring model-agent combination, so check the execution tools as well.
Insights from this ranking
- The highest score is 84.7%, achieved by GPT-5.5.
- The gap between 1st and 5th place is 4.9%p.
- OpenAI has the most models among the top 10, with 5.
- Results for this test are public for 30 models.
Autonomous computer tasks (Terminal-Bench) Epoch AI
This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04
Bars show accuracy (%) and start at 0%.
View table (top 30)
| Rank | Model | Developer | Release date | Score |
|---|---|---|---|---|
| 1 | OGPT-5.5(unknown thinking) | OpenAI | 2026-04-23 | 84.7% |
| 2 | OGPT-5.4(unknown thinking) | OpenAI | 2026-03-05 | 81.8% |
| 3 | Gemini 3.1 Pro Preview | 2026-02-19 | 80.2% | |
| 4 | Claude Opus 4.7(unknown) | Anthropic | 2026-04-16 | 80.2% |
| 5 | Claude Opus 4.6(unknown thinking) | Anthropic | 2026-02-05 | 79.8% |
| 6 | OGPT-5.3 Codex | OpenAI | 2026-02-05 | 78.4% |
| 7 | Gemini 3 Pro Preview | 2025-11-18 | 69.4% | |
| 8 | OGPT-5.2 Codex | OpenAI | 2025-12-18 | 66.5% |
| 9 | OGPT-5.2(unknown thinking) | OpenAI | 2025-12-11 | 64.9% |
| 10 | Gemini 3 Flash Preview | 2025-12-17 | 64.3% | |
| 11 | Claude Opus 4.5(unknown thinking) | Anthropic | 2025-11-24 | 63.1% |
| 12 | OGPT-5.1 Codex Mini | OpenAI | 2025-11-12 | 61.6% |
| 13 | ? | ? | 61.2% | |
| 14 | OGPT-5.1-Codex-Max | OpenAI | 2025-11-19 | 60.4% |
| 15 | OGPT-5.1 Codex | OpenAI | 2025-11-12 | 57.8% |
| 16 | xgrok-4-20 | xAI | 2026-02-17 | 57.3% |
| 17 | Claude Sonnet 4.6(unknown thinking) | Anthropic | 2026-02-17 | 53.4% |
| 18 | ZGLM-5 | Z.ai | 2026-02-11 | 52.4% |
| 19 | OGPT-5(unknown thinking) | OpenAI | 2025-08-07 | 49.6% |
| 20 | OGPT-5.1(medium) | OpenAI | 2025-11-13 | 47.6% |
| 21 | Claude Sonnet 4.5(unknown thinking) | Anthropic | 2025-09-29 | 46.5% |
| 22 | MiniMax-M2.7 | MiniMax | 2026-03-18 | 45.1% |
| 23 | OGPT-5-codex | OpenAI | 2025-09-15 | 44.3% |
| 24 | Kimi K2.5 | Moonshot AI | 2026-01-27 | 43.2% |
| 25 | MiniMax-M2.5 | MiniMax | 2026-02-12 | 42.7% |
| 26 | DeepSeek-V3.2(Thinking; Novita) | DeepSeek | 2025-12-01 | 39.6% |
| 27 | claude-opus-4-1-20250805 | Anthropic | 2025-08-05 | 38.0% |
| 28 | Claude Opus 4.1(unknown thinking) | Anthropic | 2025-08-05 | 38.0% |
| 29 | MiniMax-M2.1 | MiniMax | 2025-12-23 | 36.6% |
| 30 | Kimi K2 Thinking | Moonshot AI | 2025-11-06 | 35.7% |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarks