
Frontier models show no improvement across 30 playthroughs of the board game Earthborne Rangers, scoring far below expert humans. Epoch AI's EBR-bench probes whether AI can learn on the fly.

Give up at least one of: text only, short time horizon, easy to grade, and expert human superiority.

In this episode, Greg Burnham and Tom Adamczewski join Anson Ho to push back on benchmark pessimism and dig into what the next generation of AI benchmarks could look like.

These benchmarks track a wide range of digital work. Progress will correlate with economic utility, but tasks are too self-contained to indicate full automation.

In this episode, Daniel Litt chats with the hosts about AI’s limits in mathematics, accelerating math research, and how to measure progress on open problems.

Is this because skills generalize very well, or because developers are pushing on all benchmarks at once?

We review OSWorld, a prominent computer use benchmark. Tasks are relatively simple, many don’t require GUIs, and success often hinges on interpreting ambiguous instructions. The benchmark is also not stable over time.

57% of problems have been solved at least once.

No company has gone from $10B to $100B as fast as OpenAI projects to do.

Improved use of knowledge and precision, helpful for research, more conceptual in geometry, but limited creativity and citation issues.