AI capabilities

AI systems can now write code, pass professional exams, and assist with scientific research, and their capabilities are improving remarkably fast. But measuring exactly what AI can and cannot do is genuinely difficult, with benchmarks struggling to keep pace and real-world performance often diverging from test scores. Epoch tracks AI capabilities across tasks and benchmarks, examining how fast progress is happening, how predictable it is, and what it reveals about where the technology is heading.

Filter

Type
AI Capabilities and Benchmarking Hub
Updated Aug. 6, 2026
Score: 0.0000
AI Capabilities and Benchmarking Hub

Our database of benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks evaluated internally by Epoch AI as well as data collected from external sources. Explore trends in AI capabilities across time, by benchmark, or by model.

Disclosed CVEs: July Reached 5× the Pre-Mythos Record
Data Insight
Jul. 31, 2026
Score: 0.0000
Disclosed CVEs: July Reached 5× the Pre-Mythos Record

Notable organizations disclosed ~2,500 high- and critical-severity CVEs in July 2026, about 5× the pre-Mythos monthly record. Epoch AI's updated breakdown of vulnerability disclosures following Anthropic's Project Glasswing.

By Luke Emberson

OpenAI accidentally hacked Hugging Face — should we have seen it coming?
Newsletter
Jul. 22, 2026
Score: 0.0000
OpenAI accidentally hacked Hugging Face — should we have seen it coming?

Expert assessments and cyber benchmarks led us to expect that frontier models were capable of executing this kind of cyberattack

By Alexander Barry

AI detectors rarely flag human writing, but sometimes miss AI text imitating real authors
Data Insight
Jul. 15, 2026
Score: 0.0000
AI detectors rarely flag human writing, but sometimes miss AI text imitating real authors

We tested Pangram, GPTZero, and Originality.ai on human and AI text. All three caught nearly every passage written from a simple prompt, but missed roughly one in five passages imitating a specific author's style.

By Jaeho Lee

Disclosed CVEs: 3.5× Spike After Claude Mythos
Data Insight
Jul. 2, 2026
Score: 0.0000
Disclosed CVEs: 3.5× Spike After Claude Mythos

Notable organizations disclosed ~1,300 high- and critical-severity CVEs in June 2026, roughly 3.5× the pre-Mythos monthly record. Epoch AI's breakdown of vulnerability disclosures following Anthropic's Project Glasswing.

By Luke Emberson

GPT-4 led in ECI far longer than any other model
Data Insight
Jul. 2, 2026
Score: 0.0000
GPT-4 led in ECI far longer than any other model

GPT-4 topped Epoch's Capabilities Index (ECI) for roughly a year after its March 2023 release — no model since has led for as long. The second-longest reign, by o1, lasted barely three months.

By Jaeho Lee

Can AI Learn From Experience? EBR-Bench Results
Report
Jul. 1, 2026
Score: 0.0000
Can AI Learn From Experience? EBR-Bench Results

Frontier models show no improvement across 30 playthroughs of the board game Earthborne Rangers, scoring far below expert humans. Epoch AI's EBR-bench probes whether AI can learn on the fly.

By Benjamin Ou and Greg Burnham

MirrorCode: What's the largest software project AI can complete on its own?
Paper
Jun. 26, 2026
Score: 0.0000
MirrorCode: What's the largest software project AI can complete on its own?

MirrorCode is Epoch AI's benchmark for long-horizon coding: AI can reimplement entire programs end-to-end, with no access to the original source code.

By Tom Adamczewski, David Owen, David Rein, Florian Brand, Giles Edkins, Allen Hart, and Daniel O'Connell

Are Mythos’ cyber capabilities overhyped?
Newsletter
Jun. 11, 2026
Score: 0.0000
Are Mythos’ cyber capabilities overhyped?

Compiling all the public evidence on Mythos Preview’s cyber abilities

By Timothée Chauvin, Alexander Barry, Jean-Stanislas Denain, and Anson Ho

Open models lag state-of-the-art closed models by 4 months
Data Insight
May 29, 2026
Score: 0.0000
Open models lag state-of-the-art closed models by 4 months

Since January 2026, the most capable open-weight models have lagged frontier closed models by an average of four months, or 8 ECI points.

By Jack Edwards and Luke Emberson