Epoch AI · PhD-level science questions (GPQA Diamond)
We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.
How to read this leaderboard
What it measures: GPQA Diamond consists of 198 graduate-level multiple-choice questions in biology, chemistry and physics. Domain experts wrote and reviewed the questions, which are designed to be difficult for non-experts even with internet access.
How to read the score: The percentage (%) of correct answers. Choosing uniformly at random from 4 options gives an expected accuracy of 25%.
Caveats: As leading models approach a perfect score, this test alone may become less useful for distinguishing capabilities. It evaluates scientific knowledge and reasoning, which may differ from everyday work capabilities.
Insights from this ranking
- The highest score is 95.8%, achieved by GPT-6 Astra.
- The gap between 1st and 5th place is 1.0%p.
- OpenAI has the most models among the top 10, with 4.
- Results for this test are public for 30 models.
PhD-level science questions (GPQA Diamond) Epoch AI
This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04
Bars show accuracy (%) and start at 0%.
View table (top 30)
| Rank | Model | Developer | Release date | Score |
|---|---|---|---|---|
| 1 | #1OGPT-6 Astra(max) | OpenAI | 2026-09-03 | 95.8% |
| 2 | #2Claude Sonnet 5.5(max) | Anthropic | 2026-09-28 | 95.6% |
| 3 | #3OGPT-6.1 Sol(max) | OpenAI | 2026-09-29 | 95.4% |
| 4 | #4Gemini 3.8 Flash(high) | 2026-09-02 | 95.4% | |
| 5 | #5Gemini 3.7 Flash(high) | 2026-08-13 | 94.8% | |
| 6 | #6OGPT-5.4 Pro(xhigh) | OpenAI | 2026-03-05 | 94.6% |
| 7 | #7Gemini 3.1 Pro Preview | 2026-02-19 | 94.4% | |
| 8 | #8OGPT-6 Sol(max) | OpenAI | 2026-09-22 | 94.3% |
| 9 | #9Gemini 3.6 Flash | 2026-07-21 | 94.1% | |
| 10 | #10xGrok 4.6(high) | xAI | 2026-08-12 | 94.0% |
| 11 | #11OGPT-5.5(xhigh) | OpenAI | 2026-04-23 | 94.0% |
| 12 | #12OGPT-5.5 Pro(xhigh) | OpenAI | 2026-04-23 | 93.9% |
| 13 | #13Claude Opus 5(max) | Anthropic | 2026-07-24 | 93.9% |
| 14 | #14OGPT-5.6 Sol(max) | OpenAI | 2026-07-09 | 93.5% |
| 15 | #15xGrok 4.5(high) | xAI | 2026-07-08 | 93.4% |
| 16 | #16OGPT-5.6 Terra(max) | OpenAI | 2026-07-09 | 93.3% |
| 17 | #17OGPT-5.4(xhigh) | OpenAI | 2026-03-05 | 93.3% |
| 18 | #18Kimi K3(max) | Moonshot AI | 2026-07-16 | 93.1% |
| 19 | #19Gemini 3.5 Flash | 2026-05-19 | 92.8% | |
| 20 | #20xGrok 4.7(xhigh) | xAI | 2026-09-21 | 92.7% |
| 21 | #21Qwen3.8 Max(xhigh) | Alibaba | 2026-08-02 | 92.7% |
| 22 | #22Gemini 3 Pro Preview | 2025-11-18 | 92.6% | |
| 23 | #23Qwen3.8 Max (0902)(xhigh) | Alibaba | 2026-09-01 | 92.3% |
| 24 | #24ZGLM-5.2 | Z.ai | 2026-06-16 | 91.9% |
| 25 | #25DeepSeek V4 Pro 0813(max) | DeepSeek | 2026-08-13 | 91.7% |
| 26 | #26OGPT-5.6 Luna(max) | OpenAI | 2026-07-09 | 91.6% |
| 27 | #27OGPT-5.2(xhigh) | OpenAI | 2025-12-11 | 91.4% |
| 28 | #28Claude Opus 4.8 | Anthropic | 2026-05-28 | 91.0% |
| 29 | #29DeepSeek V4 Flash 0731(max) | DeepSeek | 2026-07-31 | 91.0% |
| 30 | #30ZGLM-5.3(max) | Z.ai | 2026-08-14 | 90.9% |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings