Valumigo

Epoch AI · Real-world professional work (GDPval)

We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.

LMArena scores come from people comparing two models' answers side by side and voting for the better one, using an Elo-based system. If score differences fall within the confidence intervals, treat the models as roughly comparable.

How to read this leaderboard

What it measures: GDPval, created by OpenAI, asks models to produce work outputs across multiple occupations. Domain experts then compare the model outputs with those of human experts anonymously.

How to read the score: The win rate is the percentage (%) of model outputs judged better than those produced by human experts. The win+tie rate is the percentage judged equal or better. Epoch's default chart currently uses the win rate.

Caveats: The evaluation relies on expert judgments and selected work tasks. Results can vary by task, so check whether the displayed metric is the win rate or the win+tie rate.

Insights from this ranking

  • The highest score is 49.7%, achieved by GPT-5.2.
  • The gap between 1st and 5th place is 9.4%p.
  • OpenAI has the most models among the top 10, with 4.
  • Results for this test are public for 11 models.

Real-world professional work (GDPval) Epoch AI

This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04

OGPT-5.2
49.7%
AnthropicClaude Opus 4.5
45.5%
Anthropicclaude-opus-4-1-20250805
43.6%
AnthropicClaude Sonnet 4.5
42.5%
GoogleGemini 3 Pro Preview
40.3%
OGPT-5
34.8%
Oo3
30.8%
Oo4-mini
25.3%
Googlegemini-2.5-pro
23.3%
xgrok-4-0709_high
21.1%
OGPT-4o
9.9%

Bars show accuracy (%) and start at 0%.

View table (top 11)
RankModelDeveloperRelease dateScore
1#1OGPT-5.2(none)OpenAI2025-12-1149.7%
2#2AnthropicClaude Opus 4.5(no thinking)Anthropic2025-11-2445.5%
3#3Anthropicclaude-opus-4-1-20250805Anthropic2025-08-0543.6%
4#4AnthropicClaude Sonnet 4.5(no thinking)Anthropic2025-09-2942.5%
5#5GoogleGemini 3 Pro PreviewGoogle2025-11-1840.3%
6#6OGPT-5(medium)OpenAI2025-08-0734.8%
7#7Oo3(medium)OpenAI2025-04-1630.8%
8#8Oo4-mini(high)OpenAI2025-04-1625.3%
9#9Googlegemini-2.5-proGoogle2025-06-0523.3%
10#10xgrok-4-0709_highxAI2025-07-0921.1%
11#11OGPT-4o(Nov 2024)OpenAI2024-11-209.9%

Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings