AI in math

Few domains test AI reasoning as clearly as mathematics, where answers can be verified automatically and the hardest problems extend to the frontier of human knowledge. Epoch tracks how AI is performing on mathematical tasks over time, including through FrontierMath, our own benchmark of expert-level problems designed to test the limits of what today's best systems can do.

Filter

Type
Claude overperforms at software engineering and underperforms at math
Data Insight
May 15, 2026
Score: 0.0000
Claude overperforms at software engineering and underperforms at math

Relative to their general Epoch Capabilities Index (ECI) values, Anthropic’s Claude models overperform on software engineering benchmarks (aggregated by the SWE-ECI) and underperform on math (Math-ECI). The SWE overperformance has been consistent across most generations, and remains in recent models. The math gap may be narrowing — Opus 4.6 and 4.7 both have Math-ECIs within 1 point of their general ECI, compared to larger gaps for earlier models.

By Alexander Barry

RIP Classic Reasoning Benchmarks. What's Next?
Newsletter
May 5, 2026
Score: 0.0000
RIP Classic Reasoning Benchmarks. What's Next?

Give up at least one of: text only, short time horizon, easy to grade, and expert human superiority.

By Greg Burnham

Are AI benchmarks doomed?
Podcast
May 1, 2026
Score: 0.0000
Are AI benchmarks doomed?

In this episode, Greg Burnham and Tom Adamczewski join Anson Ho to push back on benchmark pessimism and dig into what the next generation of AI benchmarks could look like.

By Greg Burnham, Tom Adamczewski, and Anson Ho

AI math capabilities could be jagged for a long time – Daniel Litt
Podcast
Jan. 29, 2026
Score: 0.0000
AI math capabilities could be jagged for a long time – Daniel Litt

In this episode, Daniel Litt chats with the hosts about AI’s limits in mathematics, accelerating math research, and how to measure progress on open problems.

By Daniel Litt, Greg Burnham, and Anson Ho

Benchmarking AI on unsolved math problems
Update
Jan. 27, 2026
Score: 0.0000
Benchmarking AI on unsolved math problems

Benchmarking AI on a collection of unsolved mathematics problems that have resisted serious attempts by professional mathematicians.

By The Epoch AI Team

Less than 70% of FrontierMath is within reach for today’s models
Newsletter
Oct. 17, 2025
Score: 0.0000
Less than 70% of FrontierMath is within reach for today’s models

57% of problems have been solved at least once.

By Greg Burnham

Evaluating Gemini 2.5 Deep Think's math capabilities
Report
Oct. 9, 2025
Score: 0.0000
Evaluating Gemini 2.5 Deep Think's math capabilities

Improved use of knowledge and precision, helpful for research, more conceptual in geometry, but limited creativity and citation issues.

By Greg Burnham

LLMs haven’t solved the hardest problems on math contests
Data Insight
Sep. 3, 2025
Score: 0.0000
LLMs haven’t solved the hardest problems on math contests

LLMs have come a long way on high school math contests but have yet to show that they can solve the hardest problems found on these contests.

By Greg Burnham

We didn’t learn much from the IMO
Newsletter
Aug. 7, 2025
Score: 0.0000
We didn’t learn much from the IMO

The problems gave AI only a slim chance to show new capabilities

By Greg Burnham

Quantifying the algorithmic improvement from reasoning models
Newsletter
Aug. 2, 2025
Score: 0.0000
Quantifying the algorithmic improvement from reasoning models

Reasoning models were as big of an improvement as the Transformer, at least on some benchmarks

By Anson Ho and Arden Berg