Valumigo

Epoch AI · Computer interface tasks (OSWorld 2.0)

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: XLANG Lab's OSWorld 2.0 is an evaluation consisting of 108 long-running desktop and web tasks performed in real operating systems. It is a separate, more difficult test than the original OSWorld.

How to read the score: Epoch's default metric is the full success rate (%): the percentage of tasks that pass every scoring checkpoint. This differs from partial scores for completing only some checkpoints.

Caveats: Scores vary with the execution environment, tool settings and allowed number of steps. Epoch displays results with a default budget of 500 steps, so check that conditions match.

Insights from this ranking

  • The highest score is 31.4%, achieved by Claude Opus 5.
  • The gap between 1st and 5th place is 18.4%p.
  • Anthropic has the most models among the top 9, with 4.
  • Results for this test are public for 9 models.

Computer interface tasks (OSWorld 2.0) Epoch AI

This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04

AnthropicClaude Opus 5
31.4%
OGPT-5.6 Sol
27.3%
AnthropicClaude Opus 4.8
20.6%
AnthropicClaude Opus 4.7
18.2%
OGPT-5.5
13.0%
AnthropicClaude Sonnet 4.6
9.3%
MiniMaxMiniMax-M3
4.6%
Moonshot AIKimi K2.6
4.6%
AlibabaQwen3.7 Plus
2.8%

Bars show accuracy (%) and start at 0%.

View table (top 9)
RankModelDeveloperRelease dateScore
1#1AnthropicClaude Opus 5(max)Anthropic2026-07-2431.4%
2#2OGPT-5.6 Sol(max)OpenAI2026-07-0927.3%
3#3AnthropicClaude Opus 4.8Anthropic2026-05-2820.6%
4#4AnthropicClaude Opus 4.7(max)Anthropic2026-04-1618.2%
5#5OGPT-5.5(xhigh)OpenAI2026-04-2313.0%
6#6AnthropicClaude Sonnet 4.6(medium)Anthropic2026-02-179.3%
7#7MiniMaxMiniMax-M3MiniMax2026-06-014.6%
8#8Moonshot AIKimi K2.6Moonshot AI2026-04-204.6%
9#9AlibabaQwen3.7 PlusAlibaba2026-06-022.8%

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings