Epoch AI · Advanced mathematics (FrontierMath)
We regularly fetch and display public benchmark data: LMArena user voting rankings, Epoch AI test scores, and OpenRouter usage rankings. Each leaderboard shows both its publication date and the date we retrieved it. These scores are not assigned by this site.
How to read this leaderboard
What it measures: The private Tiers 1~3 evaluation set of FrontierMath, created by Epoch AI with expert mathematicians. It ranges from advanced undergraduate to early research level and aims to reduce training data contamination by keeping evaluation problems private.
How to read the score: The percentage (%) of correct answers. Answers are submitted as Python objects in a specified format and checked automatically for correctness.
Caveats: This measures something different from routine calculation or homework assistance. Keeping tasks private does not eliminate contamination risk entirely. Check the dataset version and evaluation conditions when comparing results.
Insights from this ranking
- The highest score is 93.7%, achieved by GPT-6.1 Sol.
- The gap between 1st and 5th place is 3.9%p.
- OpenAI has the most models among the top 10, with 6.
- Results for this test are public for 30 models.
Advanced mathematics (FrontierMath) Epoch AI
This data combines scores measured by Epoch AI and published by model developers or evaluation organizations. Each test covers different models, so some may not include the latest releases yet. For models evaluated at multiple reasoning levels, we show the highest score. · Retrieved 2026-10-04
Bars show accuracy (%) and start at 0%.
View table (top 30)
| Rank | Model | Developer | Release date | Score |
|---|---|---|---|---|
| 1 | #1OGPT-6.1 Sol(max) | OpenAI | 2026-09-29 | 93.7% |
| 2 | #2OGPT-6 Astra(max) | OpenAI | 2026-09-03 | 93.7% |
| 3 | #3Claude Opus 5.5(max) | Anthropic | 2026-09-22 | 91.2% |
| 4 | #4Claude Fable 5.1(max) | Anthropic | 2026-09-01 | 90.2% |
| 5 | #5OGPT-6 Sol(max) | OpenAI | 2026-09-22 | 89.8% |
| 6 | #6OGPT-5.6 Sol(max) | OpenAI | 2026-07-09 | 89.1% |
| 7 | #7Claude Sonnet 5.5(max) | Anthropic | 2026-09-28 | 88.8% |
| 8 | #8OGPT-5.5 Pro(xhigh) | OpenAI | 2026-04-23 | 87.7% |
| 9 | #9Claude Fable 5(max) | Anthropic | 2026-06-09 | 87.0% |
| 10 | #10OGPT-5.6 Terra(max) | OpenAI | 2026-07-09 | 86.0% |
| 11 | #11Claude Opus 5(max) | Anthropic | 2026-07-24 | 85.6% |
| 12 | #12OGPT-5.5(xhigh) | OpenAI | 2026-04-23 | 85.3% |
| 13 | #13OGPT-5.4 Pro(xhigh) | OpenAI | 2026-03-05 | 82.5% |
| 14 | #14OGPT-5.6 Luna(max) | OpenAI | 2026-07-09 | 82.1% |
| 15 | #15Claude Opus 4.8 | Anthropic | 2026-05-28 | 80.0% |
| 16 | #16OGPT-6 Luna(max) | OpenAI | 2026-09-22 | 78.9% |
| 17 | #17OGPT-5.4(xhigh) | OpenAI | 2026-03-05 | 78.6% |
| 18 | #18Qwen3.8 Max(xhigh) | Alibaba | 2026-08-02 | 74.7% |
| 19 | #19Muse Spark 1.3(xhigh) | Meta | 2026-09-02 | 74.4% |
| 20 | #20OGPT-5.2 Pro | OpenAI | 2025-12-11 | 74.0% |
| 21 | #21Kimi K3(max) | Moonshot AI | 2026-07-16 | 72.2% |
| 22 | #22Gemini 3.7 Flash(high) | 2026-08-13 | 71.6% | |
| 23 | #23Claude Opus 4.7(max) | Anthropic | 2026-04-16 | 70.2% |
| 24 | #24ZGLM-5.3(max) | Z.ai | 2026-08-14 | 68.8% |
| 25 | #25Gemini 3.8 Flash(high) | 2026-09-02 | 68.4% | |
| 26 | #26OGPT-5.2(xhigh) | OpenAI | 2025-12-11 | 67.4% |
| 27 | #27xGrok 4.6(xhigh) | xAI | 2026-08-12 | 66.0% |
| 28 | #28Claude Opus 4.6(max) | Anthropic | 2026-02-05 | 66.0% |
| 29 | #29Qwen3.8 Max (0902)(xhigh) | Alibaba | 2026-09-01 | 65.6% |
| 30 | #30Claude Sonnet 5(max) | Anthropic | 2026-06-30 | 65.6% |
Data: LMArena leaderboard dataset (CC BY 4.0). Changes: Top 50 entries per leaderboard; scores are rounded. https://huggingface.co/datasets/lmarena-ai/leaderboard-datasetCompany logos are trademarks of their respective owners and are used only for identification (icons: Simple Icons).Data: Epoch AI — Capabilities & benchmarking (CC BY 4.0). https://epoch.ai/benchmarksSource: OpenRouter (openrouter.ai/rankings), as of 2026-10-04. Licensed under CC BY 4.0. https://openrouter.ai/rankings