Valumigo

GPT-6.1 Sol

OpenAI · 2026-09-29

Basic Information

Developer
OOpenAI
Release Date
2026-09-29
Overall Capability Index (ECI)
Not yet measured

Official Video

OpenAI DevDay 2026 Keynote (FULL) · OpenAI

Benchmark Scores (Epoch AI)

Each benchmark also shows the model's rank among the models evaluated.

OPhD-level science questions (GPQA Diamond) · #3 of 30
95.4%
OAdvanced mathematics (FrontierMath) · #1 of 30
93.7%
OFactual question accuracy (SimpleQA Verified) · #2 of 30
73.9%

Bars show accuracy (%) and start at 0%.

User Vote Rankings (LMArena)

Rankings based on people's votes comparing responses from two models. Models do not appear on leaderboards where they have too few votes.

  • Overall#60 of 413 · 1446
  • English#57 of 413 · 1454
  • Web development#4 of 50 · 1758
  • Image understanding (vision)#30 of 50 · 1285
  • Coding#55 of 408 · 1482
  • Creative writing#60 of 411 · 1429
  • Chinese#91 of 389 · 1467

Where to Use It

Models from OpenAI are primarily available through ChatGPT. Available models and usage limits vary by plan, so check the pricing and free access page.

View ChatGPT pricing and free access →

Hands-on Test Reports

Responding to feedback that the first round was too easy, we tested bugs across multiple files, a feature defined by a specification, and a concurrency bug in a small online store codebase. We publish all 18 run records and task files.

Claude Code vs Codex, round 2: 3 practical tasks across multiple files, each run 3 times
Scores and rankings are reproduced directly from public sources and were not measured by this site. Model names vary slightly across sources and are matched automatically, so different versions may occasionally be mixed together.

Source: Epoch AI (CC BY 4.0) · Checked 2026-10-04 · LMArena (CC BY 4.0) · Checked 2026-10-04