Greg Burnham

Greg Burnham

Greg Burnham is a researcher at Epoch AI. Prior to this, he worked at Elemental Cognition and Bridgewater Associates. He has a BA in mathematics from Princeton University.

greg@epoch.ai

Filter

Topic
Type

By Greg Burnham

Can AI Learn From Experience? EBR-Bench Results
Report
Jul. 1, 2026
Score: 0.0000
Can AI Learn From Experience? EBR-Bench Results

Frontier models show no improvement across 30 playthroughs of the board game Earthborne Rangers, scoring far below expert humans. Epoch AI's EBR-bench probes whether AI can learn on the fly.

By Benjamin Ou and Greg Burnham

RIP Classic Reasoning Benchmarks. What's Next?
Newsletter
May 5, 2026
Score: 0.0000
RIP Classic Reasoning Benchmarks. What's Next?

Give up at least one of: text only, short time horizon, easy to grade, and expert human superiority.

By Greg Burnham

Are AI benchmarks doomed?
Podcast
May 1, 2026
Score: 0.0000
Are AI benchmarks doomed?

In this episode, Greg Burnham and Tom Adamczewski join Anson Ho to push back on benchmark pessimism and dig into what the next generation of AI benchmarks could look like.

By Greg Burnham, Tom Adamczewski, and Anson Ho

What do “economic value” benchmarks tell us?
Report
Feb. 13, 2026
Score: 0.0000
What do “economic value” benchmarks tell us?

These benchmarks track a wide range of digital work. Progress will correlate with economic utility, but tasks are too self-contained to indicate full automation.

By Florian Brand and Greg Burnham

AI math capabilities could be jagged for a long time – Daniel Litt
Podcast
Jan. 29, 2026
Score: 0.0000
AI math capabilities could be jagged for a long time – Daniel Litt

In this episode, Daniel Litt chats with the hosts about AI’s limits in mathematics, accelerating math research, and how to measure progress on open problems.

By Daniel Litt, Greg Burnham, and Anson Ho

Benchmark Scores = General Capability + Claudiness
Newsletter
Nov. 20, 2025
Score: 0.0000
Benchmark Scores = General Capability + Claudiness

Is this because skills generalize very well, or because developers are pushing on all benchmarks at once?

By Greg Burnham

What does OSWorld tell us about AI’s ability to use computers?
Report
Oct. 30, 2025
Score: 0.0000
What does OSWorld tell us about AI’s ability to use computers?

We review OSWorld, a prominent computer use benchmark. Tasks are relatively simple, many don’t require GUIs, and success often hinges on interpreting ambiguous instructions. The benchmark is also not stable over time.

By Florian Brand and Greg Burnham

Less than 70% of FrontierMath is within reach for today’s models
Newsletter
Oct. 17, 2025
Score: 0.0000
Less than 70% of FrontierMath is within reach for today’s models

57% of problems have been solved at least once.

By Greg Burnham

OpenAI is projecting unprecedented revenue growth
Newsletter
Oct. 15, 2025
Score: 0.0000
OpenAI is projecting unprecedented revenue growth

No company has gone from $10B to $100B as fast as OpenAI projects to do.

By Greg Burnham

Evaluating Gemini 2.5 Deep Think's math capabilities
Report
Oct. 9, 2025
Score: 0.0000
Evaluating Gemini 2.5 Deep Think's math capabilities

Improved use of knowledge and precision, helpful for research, more conceptual in geometry, but limited creativity and citation issues.

By Greg Burnham