Tom Adamczewski

Tom Adamczewski

Tom Adamczewski started Epoch AI's benchmark engineering team. He now works on developing new evals to measure economically important AI capabilities. Before Epoch AI, he created a Monte Carlo simulation application and worked on payments technology.

tom@epoch.ai

Filter

Topic
Type

By Tom Adamczewski

MirrorCode: What's the largest software project AI can complete on its own?
Paper
Jun. 26, 2026
Score: 0.0000
MirrorCode: What's the largest software project AI can complete on its own?

MirrorCode is Epoch AI's benchmark for long-horizon coding: AI can reimplement entire programs end-to-end, with no access to the original source code.

By Tom Adamczewski, David Owen, David Rein, Florian Brand, Giles Edkins, Allen Hart, and Daniel O'Connell

Are AI benchmarks doomed?
Podcast
May 1, 2026
Score: 0.0000
Are AI benchmarks doomed?

In this episode, Greg Burnham and Tom Adamczewski join Anson Ho to push back on benchmark pessimism and dig into what the next generation of AI benchmarks could look like.

By Greg Burnham, Tom Adamczewski, and Anson Ho

MirrorCode: Evidence AI can already do some weeks-long coding tasks
Report
Apr. 10, 2026
Score: 0.0000
MirrorCode: Evidence AI can already do some weeks-long coding tasks

Early results from MirrorCode benchmark with METR: AI agents can complete weeks-long coding tasks, including reimplementing a 16,000-line codebase.

By Tom Adamczewski, David Rein, David Owen, and Florian Brand

How to run SWE-bench Verified in one hour on one machine
Update
Jul. 10, 2025
Score: 0.0000
How to run SWE-bench Verified in one hour on one machine

We are releasing a public registry of optimized Docker images for SWE-bench. This allows us to run SWE-bench Verified in 62 minutes on a single GitHub Actions VM.

By Tom Adamczewski

LLMs now accept longer inputs, and the best models can use them more effectively
Data Insight
Jun. 25, 2025
Score: 0.0000
LLMs now accept longer inputs, and the best models can use them more effectively

By Greg Burnham and Tom Adamczewski

LLM providers offer a trade-off between accuracy and speed
Data Insight
Jun. 11, 2025
Score: 0.0000
LLM providers offer a trade-off between accuracy and speed

By Greg Burnham and Tom Adamczewski

LLM responses to benchmark questions are getting longer over time
Data Insight
Apr. 17, 2025
Score: 0.0000
LLM responses to benchmark questions are getting longer over time

By Luke Emberson, Ben Cottier, Josh You, Tom Adamczewski, and Jean-Stanislas Denain

LLM inference prices have fallen rapidly but unequally across tasks
Data Insight
Mar. 12, 2025
Score: 0.0000
LLM inference prices have fallen rapidly but unequally across tasks

By Ben Cottier, Ben Snodin, David Owen, and Tom Adamczewski

A more systematic and transparent AI benchmarking hub
Update
Feb. 7, 2025
Score: 0.0000
A more systematic and transparent AI benchmarking hub

We've overhauled our AI benchmarking infrastructure to provide more transparent, systematic, and up-to-date evaluations of AI model capabilities.

By Tom Adamczewski