AI systems can now write code, pass professional exams, and assist with scientific research, and their capabilities are improving remarkably fast. But measuring exactly what AI can and cannot do is genuinely difficult, with benchmarks struggling to keep pace and real-world performance often diverging from test scores. Epoch tracks AI capabilities across tasks and benchmarks, examining how fast progress is happening, how predictable it is, and what it reveals about where the technology is heading.


Our database of benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks evaluated internally by Epoch AI as well as data collected from external sources. Explore trends in AI capabilities across time, by benchmark, or by model.

Notable organizations disclosed ~2,500 high- and critical-severity CVEs in July 2026, about 5× the pre-Mythos monthly record. Epoch AI's updated breakdown of vulnerability disclosures following Anthropic's Project Glasswing.

Expert assessments and cyber benchmarks led us to expect that frontier models were capable of executing this kind of cyberattack

We tested Pangram, GPTZero, and Originality.ai on human and AI text. All three caught nearly every passage written from a simple prompt, but missed roughly one in five passages imitating a specific author's style.

Notable organizations disclosed ~1,300 high- and critical-severity CVEs in June 2026, roughly 3.5× the pre-Mythos monthly record. Epoch AI's breakdown of vulnerability disclosures following Anthropic's Project Glasswing.

GPT-4 topped Epoch's Capabilities Index (ECI) for roughly a year after its March 2023 release — no model since has led for as long. The second-longest reign, by o1, lasted barely three months.

Frontier models show no improvement across 30 playthroughs of the board game Earthborne Rangers, scoring far below expert humans. Epoch AI's EBR-bench probes whether AI can learn on the fly.

MirrorCode is Epoch AI's benchmark for long-horizon coding: AI can reimplement entire programs end-to-end, with no access to the original source code.

Compiling all the public evidence on Mythos Preview’s cyber abilities

Since January 2026, the most capable open-weight models have lagged frontier closed models by an average of four months, or 8 ECI points.