Epoch AI · Computer interface tasks (OSWorld 2.0)
We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.
How to read this leaderboard
What it measures: XLANG Lab's OSWorld 2.0 is an evaluation consisting of 108 long-running desktop and web tasks performed in real operating systems. It is a separate, more difficult test than the original OSWorld.
How to read the score: Epoch's default metric is the full success rate (%): the percentage of tasks that pass every scoring checkpoint. This differs from partial scores for completing only some checkpoints.
Caveats: Scores vary with the execution environment, tool settings and allowed number of steps. Epoch displays results with a default budget of 500 steps, so check that conditions match.
Insights from this ranking
- The highest score is 31.4%, achieved by Claude Opus 5.
- The gap between 1st and 5th place is 18.4%p.
- Anthropic has the most models among the top 9, with 4.
- Results for this test are public for 9 models.
Computer interface tasks (OSWorld 2.0) Epoch AI
This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04
Bars show accuracy (%) and start at 0%.
View table (top 9)
| Rank | Model | Developer | Release date | Score |
|---|---|---|---|---|
| 1 | #1Claude Opus 5(max) | Anthropic | 2026-07-24 | 31.4% |
| 2 | #2OGPT-5.6 Sol(max) | OpenAI | 2026-07-09 | 27.3% |
| 3 | #3Claude Opus 4.8 | Anthropic | 2026-05-28 | 20.6% |
| 4 | #4Claude Opus 4.7(max) | Anthropic | 2026-04-16 | 18.2% |
| 5 | #5OGPT-5.5(xhigh) | OpenAI | 2026-04-23 | 13.0% |
| 6 | #6Claude Sonnet 4.6(medium) | Anthropic | 2026-02-17 | 9.3% |
| 7 | #7MiniMax-M3 | MiniMax | 2026-06-01 | 4.6% |
| 8 | #8Kimi K2.6 | Moonshot AI | 2026-04-20 | 4.6% |
| 9 | #9Qwen3.7 Plus | Alibaba | 2026-06-02 | 2.8% |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings