Few domains test AI reasoning as clearly as mathematics, where answers can be verified automatically and the hardest problems extend to the frontier of human knowledge. Epoch tracks how AI is performing on mathematical tasks over time, including through FrontierMath, our own benchmark of expert-level problems designed to test the limits of what today's best systems can do.




Relative to their general Epoch Capabilities Index (ECI) values, Anthropic’s Claude models overperform on software engineering benchmarks (aggregated by the SWE-ECI) and underperform on math (Math-ECI). The SWE overperformance has been consistent across most generations, and remains in recent models. The math gap may be narrowing — Opus 4.6 and 4.7 both have Math-ECIs within 1 point of their general ECI, compared to larger gaps for earlier models.

Give up at least one of: text only, short time horizon, easy to grade, and expert human superiority.

In this episode, Greg Burnham and Tom Adamczewski join Anson Ho to push back on benchmark pessimism and dig into what the next generation of AI benchmarks could look like.

In this episode, Daniel Litt chats with the hosts about AI’s limits in mathematics, accelerating math research, and how to measure progress on open problems.

Benchmarking AI on a collection of unsolved mathematics problems that have resisted serious attempts by professional mathematicians.

57% of problems have been solved at least once.

Improved use of knowledge and precision, helpful for research, more conceptual in geometry, but limited creativity and citation issues.

LLMs have come a long way on high school math contests but have yet to show that they can solve the hardest problems found on these contests.
The problems gave AI only a slim chance to show new capabilities
Reasoning models were as big of an improvement as the Transformer, at least on some benchmarks