Epoch AI · Factual question accuracy (SimpleQA Verified)
We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.
How to read this leaderboard
What it measures: SimpleQA Verified evaluates whether models answer short factual questions accurately. Google DeepMind and Google Research developed this 1,000-question evaluation from OpenAI's SimpleQA, reducing duplicates, answer errors and other issues.
How to read the score: Epoch displays the percentage (%) of all questions answered correctly. This differs from the F1 metric in Google's official evaluation, and this score alone cannot fully assess hallucinations or the tendency to refuse answers.
Caveats: Models answer using internal knowledge without search tools, so results may differ from the accuracy of services with search enabled. Epoch uses additional instructions to reduce answer refusals.
Insights from this ranking
- The highest score is 75.6%, achieved by GPT-6 Astra.
- The gap between 1st and 5th place is 4.8%p.
- Google has the most models among the top 10, with 4.
- Results for this test are public for 30 models.
Factual question accuracy (SimpleQA Verified) Epoch AI
This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04
Bars show accuracy (%) and start at 0%.
View table (top 30)
| Rank | Model | Developer | Release date | Score |
|---|---|---|---|---|
| 1 | #1OGPT-6 Astra(max) | OpenAI | 2026-09-03 | 75.6% |
| 2 | #2OGPT-6.1 Sol(max) | OpenAI | 2026-09-29 | 73.9% |
| 3 | #3Gemini 3.1 Pro Preview | 2026-02-19 | 73.5% | |
| 4 | #4Claude Opus 5.5(max) | Anthropic | 2026-09-22 | 72.2% |
| 5 | #5Claude Fable 5.1(max) | Anthropic | 2026-09-01 | 70.8% |
| 6 | #6Claude Fable 5(xhigh) | Anthropic | 2026-06-09 | 70.7% |
| 7 | #7Gemini 3.8 Flash(high) | 2026-09-02 | 69.7% | |
| 8 | #8OGPT-5.6 Sol(max) | OpenAI | 2026-07-09 | 69.7% |
| 9 | #9Gemini 3.7 Flash(high) | 2026-08-13 | 69.2% | |
| 10 | #10Gemini 3 Flash Preview | 2025-12-17 | 66.8% | |
| 11 | #11Gemini 3.5 Flash | 2026-05-19 | 66.2% | |
| 12 | #12Gemini 3.6 Flash | 2026-07-21 | 66.2% | |
| 13 | #13OGPT-5.5(xhigh) | OpenAI | 2026-04-23 | 63.0% |
| 14 | #14OGPT-6 Sol(max) | OpenAI | 2026-09-22 | 60.7% |
| 15 | #15Muse Spark 1.2(xhigh) | Meta | 2026-08-05 | 60.3% |
| 16 | #16Claude Opus 5(max) | Anthropic | 2026-07-24 | 59.9% |
| 17 | #17Muse Spark 1.1 | Meta | 2026-07-09 | 57.8% |
| 18 | #18xGrok 4.7(xhigh) | xAI | 2026-09-21 | 56.0% |
| 19 | #19Qwen3.7 Max | Alibaba | 2026-05-19 | 55.8% |
| 20 | #20Claude Opus 4.8 | Anthropic | 2026-05-28 | 53.0% |
| 21 | #21DeepSeek V4 Pro 0813(max) | DeepSeek | 2026-08-13 | 52.9% |
| 22 | #22Qwen 3.6 Max(Preview) | Alibaba | 2026-04-20 | 52.0% |
| 23 | #23Claude Opus 4.7(xhigh) | Anthropic | 2026-04-16 | 51.7% |
| 24 | #24Kimi K3(max) | Moonshot AI | 2026-07-16 | 50.6% |
| 25 | #25OGPT-5(high) | OpenAI | 2025-08-07 | 50.1% |
| 26 | #26Oo3(high) | OpenAI | 2025-04-16 | 49.4% |
| 27 | #27xGrok 4.6(high) | xAI | 2026-08-12 | 49.3% |
| 28 | #28Qwen3-Max-Instruct | Alibaba | 2025-09-24 | 48.7% |
| 29 | #29xGrok 4.5(high) | xAI | 2026-07-08 | 48.3% |
| 30 | #30OGPT-5.1(high) | OpenAI | 2025-11-13 | 48.0% |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings