The best score on the Epoch Capabilities Index grew almost twice as fast over the last two years as it did over the two years before that, with a 90% acceleration in April 2024. This is consistent with a similar 2024 acceleration seen in the METR Time Horizon benchmark of 40% in October of 2024. The acceleration roughly coincides with the rise of reasoning models and an increasing focus on reinforcement learning among frontier labs.
| Model | Display name | ECI | ECI CI low | ECI CI high | Date | Organization | Country (of organization) | Model accessibility | Accessibility group | Model versions |
|---|---|---|---|---|---|---|---|---|---|---|
| Gemini 3 Pro | Gemini 3 Pro Preview | 153.99 | 151.54 | 158.38 | 2025-11-18 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| Gemini 3 Flash | Gemini 3 Flash | 152.88 | 149.75 | 156.61 | 2025-12-17 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | Unknown | Other | |
| GPT-5.2 | GPT-5.2 (xhigh) | 152.03 | 148.95 | 156.95 | 2025-12-11 | OpenAI | United States of America | Unknown | Other | |
| GPT-5 Pro | GPT-5 Pro | 150.33 | 148.01 | 152.54 | 2025-10-07 | OpenAI | United States of America | API access | Closed weights | |
| GPT-5.1 | GPT-5.1 (high) | 150.25 | 148.79 | 152.00 | 2025-11-13 | OpenAI | United States of America | API access | Closed weights | |
| GPT-5 | GPT-5 (high) | 150.00 | 2025-08-07 | OpenAI | United States of America | API access | Closed weights | |||
| Claude Opus 4.5 | Claude Opus 4.5 (32k thinking) | 149.93 | 146.57 | 152.57 | 2025-11-24 | Anthropic | United States of America | API access | Closed weights | |
| o3-pro | o3-pro | 148.75 | 146.90 | 153.68 | 2025-06-10 | OpenAI | United States of America | Unknown | Other | |
| Grok 4 | Grok 4 | 147.42 | 144.63 | 149.95 | 2025-07-09 | xAI | United States of America | API access | Closed weights | |
| o3 | o3 (high) | 146.76 | 143.86 | 148.87 | 2025-04-16 | OpenAI | United States of America | API access | Closed weights | |
| Claude Sonnet 4.5 | Claude Sonnet 4.5 (59k thinking) | 146.42 | 143.06 | 148.92 | 2025-09-29 | Anthropic | United States of America | API access | Closed weights | |
| Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | 146.18 | 143.65 | 149.39 | 2025-06-17 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| Qwen3-Max | Qwen3-Max-Instruct | 145.27 | 139.50 | 152.48 | 2025-09-24 | Alibaba | China | API access | Closed weights | |
| Qwen3-235B-A22B-Thinking (Jul 2025) | Qwen3-235B-A22B-Thinking (Jul 2025) | 145.26 | 141.96 | 147.81 | 2025-07-25 | Alibaba | China | Open weights (unrestricted) | Open weights | |
| Kimi K2 Thinking | kimi-k2-thinking (turbo official) | 145.05 | 142.46 | 146.93 | 2025-11-06 | Moonshot | China | Open weights (restricted use) | Open weights | |
| o4-mini | o4-mini (high) | 144.92 | 142.15 | 146.64 | 2025-04-16 | OpenAI | United States of America | API access | Closed weights | |
| DeepSeek-V3.2-Exp | DeepSeek-V3.2-Exp | 144.66 | 141.76 | 147.61 | 2025-09-29 | DeepSeek | China | Open weights (unrestricted) | Open weights | |
| DeepSeek-V3.2 | DeepSeek-V3.2 | 144.59 | 141.39 | 146.74 | 2025-12-01 | DeepSeek | China | Unknown | Other | |
| GPT-5 mini | GPT-5 mini (high) | 144.19 | 141.33 | 146.13 | 2025-08-07 | OpenAI | United States of America | API access | Closed weights | |
| Gemini 2.5 Pro (Mar 2025) | Gemini 2.5 Pro Preview (Mar 2025) | 144.03 | 141.09 | 146.27 | 2025-04-09 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| Claude Opus 4.1 | Claude Opus 4.1 | 143.84 | 140.47 | 146.43 | 2025-08-05 | Anthropic | United States of America | API access | Closed weights | |
| Claude Opus 4 | Claude Opus 4 | 143.06 | 139.80 | 145.13 | 2025-05-22 | Anthropic | United States of America | API access | Closed weights | |
| Qwen3-Coder-480B-A35B | Qwen3-Coder-480B-A35B | 142.88 | 138.29 | 144.99 | 2025-07-31 | Alibaba | China | Open weights (unrestricted) | Open weights | |
| Gemini 2.5 Pro (May 2025) | Gemini 2.5 Pro Preview (Jun 2025) | 142.54 | 137.72 | 145.05 | 2025-05-06 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| Claude Sonnet 4 | Claude Sonnet 4 | 142.30 | 138.78 | 144.11 | 2025-05-22 | Anthropic | United States of America | API access | Closed weights | |
| o1 | o1 (high) | 142.28 | 139.85 | 143.76 | 2024-12-17 | OpenAI | United States of America | API access | Closed weights | |
| Gemini 2.5 Flash (Sep 2025) | Gemini 2.5 Flash (Sep 2025) | 142.10 | 136.30 | 144.50 | 2025-09-25 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| Gemini 2.5 Flash (May 2025) | Gemini 2.5 Flash (May 2025) | 141.91 | 138.71 | 143.29 | 2025-05-20 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| DeepSeek-R1 (May 2025) | DeepSeek-R1 (May 2025) | 141.74 | 138.37 | 143.26 | 2025-05-28 | DeepSeek | China | Open weights (unrestricted) | Open weights | |
| Claude 3.7 Sonnet | Claude 3.7 Sonnet (64k thinking) | 141.63 | 138.14 | 143.47 | 2025-02-24 | Anthropic | United States of America | API access | Closed weights | |
| Gemini 2.5 Flash (Jun 2025) | Gemini 2.5 Flash (Jun 2025) | 141.34 | 136.41 | 143.11 | 2025-06-17 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| o3-mini | o3-mini (high) | 141.00 | 135.71 | 142.76 | 2025-01-31 | OpenAI | United States of America | API access | Closed weights | |
| Grok-3 mini | Grok-3 mini | 140.70 | 137.27 | 142.75 | 2025-04-09 | xAI | United States of America | API access | Closed weights | |
| Claude Haiku 4.5 | Claude Haiku 4.5 | 140.45 | 135.99 | 142.15 | 2025-10-15 | Anthropic | United States of America | API access | Closed weights | |
| Kimi K2 | Kimi K2 Instruct | 140.36 | 136.38 | 142.39 | 2025-07-12 | Moonshot | China | Open weights (restricted use) | Open weights | |
| GPT-5 nano | GPT-5 nano (high) | 139.77 | 134.77 | 142.28 | 2025-08-07 | OpenAI | United States of America | API access | Closed weights | |
| Gemini 2.5 Flash (Apr 2025) | Gemini 2.5 Flash Preview (Apr 2025) | 139.67 | 133.80 | 141.56 | 2025-04-17 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| gpt-oss-120b | gpt-oss-120b | 139.63 | 134.51 | 142.67 | 2025-08-05 | OpenAI | United States of America | Open weights (unrestricted) | Open weights | |
| Qwen3-235B-A22B | Qwen3-235B-A22B | 139.34 | 133.55 | 141.53 | 2025-04-29 | Alibaba | China | Open weights (unrestricted) | Open weights | |
| DeepSeek-R1 | DeepSeek-R1 | 139.33 | 135.17 | 141.22 | 2025-01-20 | DeepSeek | China | Open weights (unrestricted) | Open weights | |
| Grok 3 | Grok 3 | 138.73 | 134.49 | 140.61 | 2025-04-09 | xAI | United States of America | API access | Closed weights | |
| DeepSeek-V3.1 | DeepSeek-V3.1 | 138.52 | 134.44 | 143.79 | 2025-08-21 | DeepSeek | China | Open weights (unrestricted) | Open weights | |
| GPT-4.1 | GPT-4.1 | 137.36 | 131.62 | 139.22 | 2025-04-14 | OpenAI | United States of America | API access | Closed weights | |
| GPT-4.5 | GPT-4.5 Preview (Feb 2025) | 137.16 | 131.35 | 139.25 | 2025-02-27 | OpenAI | United States of America | API access | Closed weights | |
| DeepSeek-V3 | DeepSeek-V3 (Mar 2025) | 136.50 | 131.63 | 139.16 | 2025-03-24 | DeepSeek | China | Open weights (restricted use) | Open weights | |
| o1-mini | o1-mini (high) | 136.43 | 130.83 | 138.01 | 2024-09-12 | OpenAI | United States of America | API access | Closed weights | |
| Gemini 2.0 Flash Thinking (Jan 2025) | Gemini 2.0 Flash Thinking Exp | 135.93 | 127.58 | 138.29 | 2025-01-21 | Google DeepMind,Google | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| Gemini 2.0 Pro | Gemini 2.0 Pro Exp (Feb 2025) | 135.34 | 128.91 | 137.51 | 2025-02-05 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | Hosted access (no API) | Closed weights | |
| GPT-4.1 mini | GPT-4.1 mini | 135.27 | 128.37 | 137.34 | 2025-04-14 | OpenAI | United States of America | API access | Closed weights | |
| o1-preview | o1-preview | 135.24 | 129.71 | 139.67 | 2024-09-12 | OpenAI | United States of America | API access | Closed weights | |
| Mistral Medium 3 | Mistral Medium 3 | 134.88 | 125.62 | 136.49 | 2025-05-07 | Mistral AI | France | API access | Closed weights | |
| Gemini 2.0 Flash | Gemini 2.0 Flash (Feb 2025) | 134.86 | 128.12 | 136.96 | 2025-02-05 | Google DeepMind,Google | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| Claude 3.5 Sonnet (October 2024) | Claude 3.5 Sonnet (Oct 2024) | 134.10 | 128.96 | 139.08 | 2024-10-22 | Anthropic | United States of America | API access | Closed weights | |
| Qwen2.5-Max | Qwen2.5-Max | 132.97 | 125.09 | 136.40 | 2025-01-25 | Alibaba | China | API access | Closed weights | |
| Llama 4 Maverick | Llama 4 Maverick (FP8) | 132.69 | 122.84 | 134.71 | 2025-04-05 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| Gemini 1.5 Pro | Gemini 1.5 Pro | 132.32 | 125.18 | 134.01 | 2024-05-24 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| Qwen Plus | Qwen Plus | 130.83 | 122.08 | 132.76 | 2025-04-28 | Alibaba | China | API access | Closed weights | |
| Phi-4 | Phi-4 | 130.76 | 121.84 | 132.95 | 2024-12-12 | Microsoft Research | United States of America | Open weights (unrestricted) | Open weights | |
| Gemma 3 27B | Gemma 3 27B | 130.72 | 120.11 | 133.28 | 2025-03-12 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | Open weights (restricted use) | Open weights | |
| GPT-4.1 nano | GPT-4.1 nano | 130.50 | 121.74 | 133.32 | 2025-04-14 | OpenAI | United States of America | API access | Closed weights | |
| Grok-2 | Grok-2 | 130.37 | 121.99 | 131.93 | 2024-12-12 | xAI | United States of America | API access | Closed weights | |
| Gemini 1.5 Flash (Sep 2024) | Gemini 1.5 Flash (Sep 2024) | 130.00 | 118.63 | 132.21 | 2024-09-24 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| Claude 3.5 Sonnet | Claude 3.5 Sonnet (Jun 2024) | 130.00 | 2024-06-20 | Anthropic | United States of America | API access | Closed weights | |||
| Llama 4 Scout | Llama 4 Scout | 129.83 | 120.98 | 131.64 | 2025-04-05 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| GPT-4o (Nov 2024) | GPT-4o (Nov 2024) | 129.81 | 120.58 | 132.98 | 2024-11-20 | OpenAI | United States of America | API access | Closed weights | |
| Qwen2.5-72B | Qwen2.5-72B | 129.26 | 119.80 | 131.59 | 2024-09-19 | Alibaba | China | Open weights (unrestricted) | Open weights | |
| GPT-4o (Aug 2024) | GPT-4o (Aug 2024) | 128.94 | 119.42 | 131.74 | 2024-08-06 | OpenAI | United States of America | API access | Closed weights | |
| GPT-4o (May 2024) | GPT-4o (May 2024) | 128.33 | 117.87 | 130.63 | 2024-05-13 | OpenAI | United States of America | API access | Closed weights | |
| Mistral Large 2 | Mistral Large 2 | 128.25 | 119.47 | 130.50 | 2024-11-18 | Mistral AI | France | Open weights (non-commercial) | Open weights | |
| Llama 3.1-405B | Llama 3.1-405B | 127.72 | 114.66 | 131.02 | 2024-07-23 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| Claude 3.5 Haiku | Claude 3.5 Haiku (Oct 2024) | 127.33 | 115.28 | 130.98 | 2024-10-22 | Anthropic | United States of America | API access | Closed weights | |
| GPT-4 Turbo (Apr 2024) | GPT-4 Turbo (Apr 2024) | 127.25 | 116.69 | 129.51 | 2024-04-09 | OpenAI | United States of America | API access | Closed weights | |
| GPT-4o mini | GPT-4o mini | 127.00 | 118.62 | 129.74 | 2024-07-18 | OpenAI | United States of America | API access | Closed weights | |
| Llama 3.3 70B | Llama 3.3 70B | 126.93 | 116.98 | 129.77 | 2024-12-06 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| Claude 3 Opus | Claude 3 Opus | 126.57 | 118.27 | 130.99 | 2024-02-29 | Anthropic | United States of America | API access | Closed weights | |
| GPT-4 (Mar 2023) | GPT-4 (Mar 2023) | 125.80 | 114.60 | 132.66 | 2023-03-14 | OpenAI | United States of America | API access | Closed weights | |
| DeepSeek-V2 (MoE-236B) | DeepSeek-V2 (MoE-236B) | 125.49 | 115.91 | 128.75 | 2024-05-07 | DeepSeek | China | Open weights (restricted use) | Open weights | |
| Qwen2-72B | Qwen2-72B | 125.33 | 111.98 | 127.72 | 2024-06-07 | Alibaba | China | Open weights (unrestricted) | Open weights | |
| Llama 3.2 90B | Llama 3.2 90B | 125.30 | 115.22 | 127.83 | 2024-09-24 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| Llama 3.1-70B | Llama 3.1-70B | 125.07 | 114.47 | 128.08 | 2024-07-23 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| Gemini 1.5 Flash (May 2024) | Gemini 1.5 Flash (May 2024) | 122.47 | 110.06 | 125.10 | 2024-05-23 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| phi-3-small 7.4B | phi-3-small 7.4B | 122.29 | 107.47 | 126.63 | 2024-04-23 | Microsoft | United States of America | Open weights (unrestricted) | Open weights | |
| Gemma 2 27B | Gemma 2 27B | 122.20 | 109.06 | 125.04 | 2024-06-24 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | Open weights (restricted use) | Open weights | |
| GPT-4 (Jun 2023) | GPT-4 (Jun 2023) | 121.85 | 109.32 | 124.50 | 2023-06-13 | OpenAI | United States of America | API access | Closed weights | |
| Claude Instant | Claude Instant | 121.45 | 108.26 | 125.53 | Anthropic | United States of America | API access | Closed weights | ||
| Qwen2.5-Coder (32B) | Qwen2.5-Coder (32B) | 121.23 | 103.05 | 127.70 | 2024-09-18 | Alibaba | China | Open weights (unrestricted) | Open weights | |
| phi-3-medium 14B | phi-3-medium 14B | 121.01 | 108.96 | 125.45 | 2024-04-23 | Microsoft | United States of America | Open weights (unrestricted) | Open weights | |
| Mixtral 8x22B | Mixtral 8x22B | 120.93 | 106.44 | 123.60 | 2024-04-17 | Mistral AI | France | Open weights (unrestricted) | Open weights | |
| Mistral Large | Mistral Large | 120.88 | 107.91 | 124.40 | 2024-02-26 | Mistral AI | France | API access | Closed weights | |
| Claude 3 Sonnet | Claude 3 Sonnet | 119.71 | 105.64 | 123.26 | 2024-02-29 | Anthropic | United States of America | API access | Closed weights | |
| Gemma 2 9B | Gemma 2 9B | 119.53 | 103.49 | 122.69 | 2024-06-24 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | Open weights (restricted use) | Open weights | |
| Claude 2 | Claude 2 | 119.16 | 106.05 | 125.50 | 2023-07-11 | Anthropic | United States of America | API access | Closed weights | |
| Llama 3-70B | Llama 3-70B | 119.12 | 93.97 | 125.08 | 2024-04-18 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| GPT-3.5 | GPT-3.5 | 118.62 | 104.25 | 122.16 | 2022-03-15 | OpenAI | United States of America | API access | Closed weights | |
| Mixtral 8x7B | Mixtral 8x7B | 118.23 | 104.76 | 122.09 | 2023-12-11 | Mistral AI | France | Open weights (unrestricted) | Open weights | |
| Mistral NeMo | Mistral NeMo | 118.12 | 104.62 | 122.49 | 2024-07-18 | Mistral AI | France | Open weights (unrestricted) | Open weights | |
| Claude 2.1 | Claude 2.1 | 117.83 | 103.60 | 121.16 | 2023-11-21 | Anthropic | United States of America | API access | Closed weights | |
| Yi-34B | Yi-34B | 117.47 | 101.53 | 121.73 | 2023-11-22 | 01.AI | China | Open weights (restricted use) | Open weights | |
| phi-3-mini 3.8B | phi-3-mini 3.8B | 117.25 | 103.94 | 122.56 | 2024-04-23 | Microsoft | United States of America | Open weights (unrestricted) | Open weights | |
| GPT-3.5 Turbo | GPT-3.5 Turbo (Nov 2023) | 117.15 | 100.55 | 121.93 | 2023-11-06 | OpenAI | United States of America | API access | Closed weights | |
| Gemini 1.0 Pro | Gemini 1.0 Pro | 116.65 | 101.51 | 120.79 | 2024-02-15 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | Closed weights | |
| Stable Beluga 2 | Stable Beluga 2 | 116.59 | 99.95 | 120.24 | 2023-07-20 | Stability AI | United Kingdom of Great Britain and Northern Ireland | Open weights (non-commercial) | Open weights | |
| Llama 3-8B | Llama 3-8B | 116.57 | 102.21 | 119.68 | 2024-04-18 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| Claude 3 Haiku | Claude 3 Haiku | 116.16 | 98.18 | 121.75 | 2024-03-07 | Anthropic | United States of America | API access | Closed weights | |
| Llama 3.1-8B | Llama 3.1-8B | 115.45 | 94.32 | 121.44 | 2024-07-23 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| PaLM (540B) | PaLM (540B) | 114.63 | 98.92 | 118.78 | 2022-04-04 | Google Research | United States of America | Unreleased | Closed weights | |
| Llama 2-70B | Llama 2-70B | 113.64 | 96.24 | 117.38 | 2023-07-18 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| Qwen-14B | Qwen-14B | 112.04 | 92.64 | 117.25 | 2023-09-28 | Alibaba | China | Open weights (restricted use) | Open weights | |
| Qwen2.5-Coder (7B) | Qwen2.5-Coder (7B) | 112.02 | 93.48 | 119.83 | 2024-09-18 | Alibaba | China | Open weights (unrestricted) | Open weights | |
| Falcon-180B | Falcon-180B | 111.30 | 95.34 | 119.88 | 2023-09-06 | Technology Innovation Institute | United Arab Emirates | Open weights (restricted use) | Open weights | |
| Mistral 7B | Mistral 7B | 111.18 | 93.94 | 116.67 | 2024-05-27 | Mistral AI | France | Open weights (unrestricted) | Open weights | |
| Gemma 7B | Gemma 7B | 110.98 | 94.05 | 116.65 | 2024-02-21 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | Open weights (restricted use) | Open weights | |
| Chinchilla | Chinchilla | 110.95 | 89.86 | 116.42 | 2022-03-29 | DeepMind | United Kingdom of Great Britain and Northern Ireland | Unreleased | Closed weights | |
| Megatron-Turing NLG 530B | Megatron-Turing NLG 530B | 109.97 | 93.94 | 127.20 | 2022-01-28 | Microsoft,NVIDIA | United States of America | Unreleased | Closed weights | |
| GLaM | GLaM | 109.02 | 89.12 | 116.20 | 2021-12-13 | United States of America | Unreleased | Closed weights | ||
| LLaMA-65B | LLaMA-65B | 108.98 | 89.84 | 113.57 | 2023-02-24 | Meta AI | United States of America | Open weights (non-commercial) | Open weights | |
| Falcon 2 11B | Falcon 2-11B | 107.97 | 88.50 | 117.23 | 2024-05-09 | Technology Innovation Institute | United Arab Emirates | Open weights (restricted use) | Open weights | |
| Phi-2 | Phi-2 | 105.99 | 79.12 | 112.92 | 2023-12-12 | Microsoft | United States of America | Open weights (unrestricted) | Open weights | |
| LLaMA-33B | LLaMA-33B | 105.62 | 84.76 | 111.68 | 2023-02-24 | Meta AI | United States of America | Open weights (non-commercial) | Open weights | |
| Nemotron-4 15B | Nemotron-4 15B | 105.43 | 85.62 | 111.65 | 2024-02-26 | NVIDIA | United States of America | Unreleased | Closed weights | |
| Qwen-7B | Qwen-7B | 104.96 | 76.24 | 112.78 | 2023-09-28 | Alibaba | China | Open weights (restricted use) | Open weights | |
| InstructGPT 175B | InstructGPT 175B | 104.57 | 79.98 | 111.67 | 2022-01-27 | OpenAI | United States of America | API access | Closed weights | |
| Gopher (280B) | Gopher (280B) | 104.03 | 82.31 | 109.54 | 2021-12-08 | DeepMind | United Kingdom of Great Britain and Northern Ireland | Unreleased | Closed weights | |
| Llama 2-13B | Llama 2-13B | 103.38 | 78.85 | 110.08 | 2023-07-18 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| Llama 2-34B | Llama 2-34B | 102.91 | 80.18 | 109.89 | 2023-07-18 | Meta AI | United States of America | Unreleased | Closed weights | |
| StarCoder 2 15B | StarCoder 2 15B | 102.69 | 81.11 | 111.57 | 2024-02-20 | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | Open weights (restricted use) | Open weights | |
| Yi 6B | Yi 6B | 102.20 | 82.85 | 109.51 | 2023-11-02 | 01.AI | China | Open weights (restricted use) | Open weights | |
| Falcon-40B | Falcon-40B | 101.94 | 79.13 | 108.72 | 2023-03-15 | Technology Innovation Institute | United Arab Emirates | Open weights (unrestricted) | Open weights | |
| Qwen2.5-Coder (1.5B) | Qwen2.5-Coder (1.5B) | 100.38 | 70.93 | 109.84 | 2024-09-18 | Alibaba | China | Open weights (unrestricted) | Open weights | |
| Baichuan2-13B | Baichuan2-13B | 100.27 | 75.99 | 110.32 | 2023-09-06 | Baichuan | China | Open weights (restricted use) | Open weights | |
| MPT-30B | MPT-30B | 97.39 | 70.25 | 104.83 | 2023-06-22 | MosaicML | United States of America | Open weights (unrestricted) | Open weights | |
| LLaMA-13B | LLaMA-13B | 96.47 | 73.50 | 104.80 | 2023-02-24 | Meta AI | United States of America | Open weights (non-commercial) | Open weights | |
| Llama 2-7B | Llama 2-7B | 94.49 | 68.27 | 103.44 | 2023-07-18 | Meta AI | United States of America | Open weights (restricted use) | Open weights | |
| DeepSeek Coder 33B | DeepSeek Coder 33B | 92.74 | 64.80 | 105.02 | 2023-11-02 | DeepSeek,Peking University | China | Open weights (restricted use) | Open weights | |
| Baichuan 2-7B | Baichuan 2-7B | 92.41 | 62.00 | 102.31 | 2023-09-20 | Baichuan | China | Open weights (restricted use) | Open weights | |
| LLaMA-7B | LLaMA-7B | 91.44 | 65.58 | 100.61 | 2023-02-24 | Meta AI | United States of America | Open weights (non-commercial) | Open weights | |
| Falcon-7B | Falcon-7B | 89.88 | 63.05 | 100.02 | 2023-04-24 | Technology Innovation Institute | United Arab Emirates | Open weights (unrestricted) | Open weights | |
| Gemma 2B | Gemma 2B | 89.79 | 61.83 | 99.54 | 2024-02-21 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | Open weights (restricted use) | Open weights | |
| StarCoder 2 7B | StarCoder 2 7B | 89.45 | 58.31 | 100.21 | 2024-02-20 | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | Open weights (restricted use) | Open weights | |
| MPT-7B | MPT-7B | 89.18 | 60.64 | 97.57 | 2023-05-05 | MosaicML | United States of America | Open weights (unrestricted) | Open weights | |
| GPT-NeoX-20B | GPT-NeoX-20B | 87.40 | 63.85 | 99.58 | 2022-04-07 | EleutherAI | United States of America | Open weights (unrestricted) | Open weights | |
| XGen-7B | XGen-7B | 87.31 | 58.08 | 97.03 | 2023-06-27 | Salesforce | United States of America | Open weights (unrestricted) | Open weights | |
| Baichuan1-7B | Baichuan1-7B | 85.72 | 55.41 | 96.73 | 2023-06-01 | Baichuan | China | Open weights (non-commercial) | Open weights | |
| Phi-1.5 | Phi-1.5 | 84.89 | 33.48 | 100.07 | 2023-09-11 | Microsoft | United States of America | Open weights (unrestricted) | Open weights | |
| DeepSeek Coder 6.7B | DeepSeek Coder 6.7B | 84.74 | 53.75 | 95.49 | 2023-11-02 | DeepSeek,Peking University | China | Open weights (restricted use) | Open weights | |
| StarCoder 2 3B | StarCoder 2 3B | 83.64 | 45.52 | 95.94 | 2024-02-22 | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | Open weights (restricted use) | Open weights | |
| Dolly 2.0-12b | Dolly 2.0-12b | 81.86 | 51.00 | 94.81 | 2023-04-11 | Databricks | United States of America | Open weights (unrestricted) | Open weights | |
| Cerebras-GPT-13B | Cerebras-GPT-13B | 72.92 | 39.95 | 89.32 | 2023-03-20 | Cerebras Systems | United States of America | Open weights (unrestricted) | Open weights | |
| DeepSeek Coder 1.3B | DeepSeek Coder 1.3B | 51.73 | 5.27 | 72.72 | 2023-11-02 | DeepSeek,Peking University | China | Open weights (restricted use) | Open weights | |
| OPT-1.3B | OPT-1.3B | 50.32 | 10.23 | 119.27 | 2022-05-11 | Meta AI | United States of America | Open weights (non-commercial) | Open weights |
| Model version | Time horizon | Release date | Organization | Country | Training compute (FLOP) | Training compute notes | CI high | CI low | Source | Source link | Average score | Notes | ID |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| claude-3-5-sonnet-20240620 | 18.68 | 2024-06-20 | Anthropic | United States of America | 2.7e+25 | Blog post by Dario Amodei includes some info on 3.5 Sonnet compute: https://darioamodei.com/on-deepseek-and-export-controls
"Claude 3.5 Sonnet is a mid-sized model that cost a few $10M's to train (I won't give an exact number). Also, 3.5 Sonnet was not trained in any way that involved a larger or more expensive model (contrary to some rumors)."
Using assumptions about GPU pricing, this lets us estimate compute. https://docs.google.com/spreadsheets/d/1-p-ab6t6dkUM6T7GwnFp85ePTMpZMW7LFY7fW2t8POs/ | 34.39 | 9.47 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.48 | rec3I5CcPGjF3Riat | |
| claude-3-5-sonnet-20241022 | 29.58 | 2024-10-22 | Anthropic | United States of America | 59.06 | 14.03 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.53 | recM2JJhzA3vJAZoZ | |||
| claude-3-7-sonnet-20250219_16K | 56.09 | 2025-02-24 | Anthropic | United States of America | 3.4e+25 | https://docs.google.com/spreadsheets/d/10bhwdVrfHI8tysVIz62ZxtvQ30L-HojYvmU18_b-WIM/edit?gid=0#gid=0 | 93.42 | 29.33 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.60 | recqRojHxRU5ySXWq | |
| claude-3-opus-20240229 | 6.41 | 2024-02-29 | Anthropic | United States of America | Training compute estimated to be 1.64e25 FLOP from benchmark scores. https://colab.research.google.com/drive/1r3pUMhB7Kh0Gls9eG-v_XefWrye9fVQR?usp=sharing | 12.90 | 2.91 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.38 | recyvUC4eTQ5SK6Gr | ||
| claude-opus-4-1-20250805_16K | 113.69 | 2025-08-05 | Anthropic | United States of America | 209.19 | 57.55 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.67 | recuz9bui4PgiPuGH | |||
| claude-opus-4-5-20251101_16K | 288.90 | 2025-11-24 | Anthropic | United States of America | Flagship model from a leading developer in mid-2025; very likely it used >1e25 FLOP. | 1,225.19 | 108.87 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.75 | recPjuZumaNjbylqU | ||
| claude-opus-4-20250514_16K | 85.56 | 2025-05-22 | Anthropic | United States of America | Flagship model from a leading developer in mid-2025; very likely it used >1e25 FLOP. | 146.80 | 45.06 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.64 | recjchvjch6vUVmzs | ||
| claude-sonnet-4-5-20250929_16K | 121.95 | 2025-09-29 | Anthropic | United States of America | 254.33 | 56.04 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.67 | recwCLGT9Kr3iDcqu | |||
| claude-sonnet-4-20250514_16K | 74.94 | 2025-05-22 | Anthropic | United States of America | Flagship model from a leading developer in mid-2025; very likely it used >1e25 FLOP. | 133.99 | 37.85 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.62 | recueBmSBs2QlEWPm | ||
| davinci-002 | 0.15 | 2020-05-28 | OpenAI | United States of America | 3.1e+23 | Table D.1 https://arxiv.org/abs/2005.14165 | 0.25 | 0.07 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.16 | recbmHH5GLQL9jfZn | |
| DeepSeek-R1 | 26.93 | 2025-01-20 | DeepSeek | China | 4.0e+24 | Estimates by Ege Erdil in Gradient Updates:
https://epoch.ai/gradient-updates/what-went-into-training-deepseek-r1
"A dataset size of 14.8 trillion tokens is reasonable and in line with other models of this scale. Assuming that’s valid, the pretraining of this model would have required 6 * (37 billion) * (14.8 trillion) = 3e24 FLOP. If we assume DeepSeek’s training cluster consists of H800s with the PCIe form factor, then each should be capable of 1.5e15 FP8 per second, and the implied model FLOP utilization (MFU) of DeepSeek v3’s 55 day training run ends up being around 23%."
6 FLOP/token/param * 14.8T tokens * 37B active params = 3.29e24 FLOP (pretraining)
1.2e23 FLOP (post-training)
6.1e23 FLOP (fine-tuning)
Total compute: 3.29e24 + 1.2e23 + 6.1e23 = 4.02e24 | 50.98 | 13.48 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.52 | rectWm0PAizmHcDik | |
| DeepSeek-R1-0528 | 31.17 | 2025-05-28 | DeepSeek | China | 4.0e+24 | Estimates by Ege Erdil in Gradient Updates:
https://epoch.ai/gradient-updates/what-went-into-training-deepseek-r1
"A dataset size of 14.8 trillion tokens is reasonable and in line with other models of this scale. Assuming that’s valid, the pretraining of this model would have required 6 * (37 billion) * (14.8 trillion) = 3e24 FLOP. If we assume DeepSeek’s training cluster consists of H800s with the PCIe form factor, then each should be capable of 1.5e15 FP8 per second, and the implied model FLOP utilization (MFU) of DeepSeek v3’s 55 day training run ends up being around 23%."
6 FLOP/token/param * 14.8T tokens * 37B active params = 3.29e24 FLOP (pretraining)
1.2e23 FLOP (post-training)
6.1e23 FLOP (fine-tuning)
Total compute: 3.29e24 + 1.2e23 + 6.1e23 = 4.02e24 | 64.11 | 12.85 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.54 | recasKfxLMpd8pTjF | |
| DeepSeek-V3 | 18.47 | 2024-12-26 | DeepSeek | China | 3.4e+24 | "At an economical cost of only 2.664M H800 GPU hours, we complete the pre-training of DeepSeek-V3 on 14.8T tokens, producing the currently strongest open-source base model. The subsequent training stages after pre-training require only 0.1M GPU hours."
6 * 37B (active params) * 14.8T = 3.2856e24 for pretraining.
We know they trained in FP8. H800s get 1.513e15 FLOP/s in FP8:
2.688M * 3600 * 1.513e15 * MFU = 3.2856e24
Suggests a MFU of 0.2244 in pre-training. If we assume MFU was the same in post-training, that adds an additional:
0.1M * 3600 * 1.513e15 * 0.2244 = 1.222e23 FLOP from post-training
Total: 3.2856e24 + 1.222e23 = 3.4078e24 FLOP | 33.77 | 8.97 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.47 | rec3j0U1Ecav7Q1FA | |
| DeepSeek-V3-0324 | 23.12 | 2025-03-24 | DeepSeek | China | 3.4e+24 | "At an economical cost of only 2.664M H800 GPU hours, we complete the pre-training of DeepSeek-V3 on 14.8T tokens, producing the currently strongest open-source base model. The subsequent training stages after pre-training require only 0.1M GPU hours."
6 * 37B (active params) * 14.8T = 3.2856e24 for pretraining.
We know they trained in FP8. H800s get 1.513e15 FLOP/s in FP8:
2.688M * 3600 * 1.513e15 * MFU = 3.2856e24
Suggests a MFU of 0.2244 in pre-training. If we assume MFU was the same in post-training, that adds an additional:
0.1M * 3600 * 1.513e15 * 0.2244 = 1.222e23 FLOP from post-training
Total: 3.2856e24 + 1.222e23 = 3.4078e24 FLOP | 41.33 | 11.71 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.50 | rec5tXt66EEIgBcLi | |
| gemini-2.5-pro-preview-06-05 | 38.73 | 2025-06-05 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | Flagship model from a leading developer in mid-2025; very likely it used >1e25 FLOP. | 69.56 | 19.32 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.55 | recxHthxDmMulkuth | ||
| gpt-3.5-turbo-instruct | 0.60 | 2023-09-18 | OpenAI | United States of America | 0.99 | 0.23 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.21 | rec7jJ3gfq0bIgHN4 | |||
| gpt-4-0125-preview | 5.37 | 2024-01-25 | OpenAI | United States of America | Training compute estimated to be 2.2e25 FLOP using benchmark imputation. https://colab.research.google.com/drive/1r3pUMhB7Kh0Gls9eG-v_XefWrye9fVQR?usp=sharing | 9.76 | 2.80 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.35 | recD86WJin9bt0ueq | ||
| gpt-4-0314 | 5.36 | 2023-03-14 | OpenAI | United States of America | 2.1e+25 | 90% CI: 8.2E+24 to 4.4E+25
NOTE: this is a rough estimate based on public information, much less information than most other systems in the database.
Calculation and confidence intervals here: https://colab.research.google.com/drive/1O99z9b1I5O66bT78r9ScslE_nOj5irN9?usp=sharing | 9.73 | 2.45 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.36 | recO1ntpu18mX9GIG | |
| gpt-4-1106-preview | 8.56 | 2023-11-06 | OpenAI | United States of America | Training compute estimated to be 2.2e25 FLOP using benchmark imputation. https://colab.research.google.com/drive/1r3pUMhB7Kh0Gls9eG-v_XefWrye9fVQR?usp=sharing | 16.08 | 4.16 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.40 | rechSoSit7zTyOpq0 | ||
| gpt-4-turbo-2024-04-09 | 6.57 | 2024-04-09 | OpenAI | United States of America | Training compute estimated to be 2.2e25 FLOP using benchmark imputation. https://colab.research.google.com/drive/1r3pUMhB7Kh0Gls9eG-v_XefWrye9fVQR?usp=sharing | 12.20 | 3.22 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.37 | recen58457Be84Pud | ||
| gpt-4o-2024-11-20 | 9.17 | 2024-11-20 | OpenAI | United States of America | Training compute estimated to be 3.8e25 FLOP from benchmark scores. https://colab.research.google.com/drive/1r3pUMhB7Kh0Gls9eG-v_XefWrye9fVQR?usp=sharing | 17.96 | 4.22 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.41 | recdwVn3QiDENdt6P | ||
| gpt-5-2025-08-07_medium | 137.32 | 2025-08-07 | OpenAI | United States of America | 6.6e+25 | Likely around 6e25 [CI: 2e25 to 2e26] FLOP. See document below for details
https://docs.google.com/document/d/1V2jIk365LnhH4WDoCw5dYJjZr1Htw8IHaK1noMf5Y48/edit?tab=t.z871imftkus | 271.12 | 66.87 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.70 | recNwo7dbEyB1crye | |
| gpt-5.1-codex-max | 161.75 | 2025-11-19 | OpenAI | United States of America | 348.27 | 74.81 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.72 | recio8mNX4NbLwCsZ | |||
| gpt-oss-120b | 41.98 | 2025-08-05 | OpenAI | United States of America | 4.9e+24 | "The training run for gpt-oss-120b required 2.1 million H100-hours to complete"
(2.1e6 hours)*(1,979 H100 FLOP/s)*(30% utilization)*(60*60) = 4.49e24
They also do post training similar to o3, which we assume adds at least 10% as much compute, so we multiply this estimate by 1.1 to get 4.94e24
| 81.51 | 18.79 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.57 | rec2VBZIfOKcr9NeL | |
| gpt2-xl | 0.04 | 2019-11-05 | OpenAI | United States of America | 1.9e+21 | Estimating based on compute = 6 FLOP/token/param * epochs * parameters * tokens.
40GB dataset is approximately 8B words, or 1/0.75 * 8B = 10.66B tokens.
The number of epochs is not reported, but another paper [1] claims in table 1 that it is 20 or 100 epochs, and another paper [2] claims 12 epochs based on communication with the GPT-2 authors (page 4).
12 epochs is the modal, most credible value. Mean of probability mass is probably around 20 epochs, so calculating from that value:
6 * (40 * 200 million * 1/0.75 * 20) * 1.5 billion parameters = 1.92e21
https://www.wolframalpha.com/input?i=6+FLOP+*+20+*+%2840+billion+%2F+5+*+%284%2F3%29%29+*+1.5+billion
[1] https://arxiv.org/abs/1906.06669 One Epoch Is All You Need
[2] https://www.usenix.org/system/files/sec21-carlini-extracting.pdf Extracting Data From Large Language Models
It also appears the model was trained on TPU v3 chips:
https://huggingface.co/openai-community/gpt2 | 0.13 | 0.00 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.10 | rec8JOj4dfN5bKICN | |
| grok-4-0709 | 110.08 | 2025-07-09 | xAI | United States of America | 5.0e+26 | We think that RL relative to pre-compute is between our estimate for o3 (10% of pre-training) and the 100% implied by this slide in the launch ( https://archive.is/f0vJU ). Assuming the same pre-training as Grok 3 (also implied by that slide, and much more consistent) and that Grok 3 used a tenth as much RL, we get:
2 * (grok3/1.1) in the high case (rl is 10% of grok 3, so grok3/1.1 is grok3 precompute, and in this case twice that is grok 4)
1.1 * (grok3/1.01) in the low case
The geometric mean is (rounded to one sig fig): 5e26
| 231.84 | 48.19 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.67 | recBhYPAeAQhBoGGs | |
| kimi-k2-thinking | 54.19 | 2025-11-06 | Moonshot | China | 4.2e+24 | Assuming the additional post-training contributed between 1% and 100% of Kimi K2's training compute (which we confidently estimate at 2.976e+24), we get a range of 3.0e24 to 6.0e24, and a geometric mean of 4.2e24. | 98.25 | 25.48 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.59 | recmuexxjm6ucxXtH | |
| o1-2024-12-17_medium | 39.21 | 2024-12-17 | OpenAI | United States of America | 84.35 | 17.60 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.56 | recrH7UpKH2XmiFW8 | |||
| o1-preview-2024-09-12 | 22.22 | 2024-09-12 | OpenAI | United States of America | 40.79 | 11.62 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.49 | recJhENih6N50V43Q | |||
| o3-2025-04-16_medium | 91.27 | 2025-04-16 | OpenAI | United States of America | 163.13 | 45.63 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.65 | recFSX9NsirVMms25 | |||
| o4-mini-2025-04-16_medium | 76.51 | 2025-04-16 | OpenAI | United States of America | We can’t make a precise estimate, but seems unlikely to exceed 10^25 FLOP. We think active parameter count is 10-30B. This would require >55T tokens to reach 10^25 FLOP at the large size, i.e. well beyond 10x overtraining relative to Chinchilla. | 151.48 | 35.21 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.64 | recTsdWFhES9SchXH | ||
| qwen2-72b-instruct | 2.24 | 2024-06-07 | Alibaba | China | 3.0e+24 | 72 billion params, 7 trillion tokens
6 * 72 billion * 7 trillion ~= 3.02e24 | 4.83 | 0.85 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.30 | recIqQ9tqJzeP1sHf | |
| qwen2.5-72b-instruct | 5.17 | 2024-09-19 | Alibaba | China | 7.8e+24 | Training dataset size was 18 trillion
6ND = 6 * 72.7 billion parameters * 18 trillion tokens = 7.8e24 | 10.33 | 2.38 | METR - Measuring AI Ability to Complete Long Tasks | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ | 0.36 | recmXpRGOYoAcTAgQ |
In order to provide additional data we include pre-2023 models, which are currently filtered out from the Benchmarking Hub due to the comparative sparsity of benchmark scores during that time. Our conclusions do not change substantially if this data is excluded.
Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.
Learn more about this graph
Using the Epoch Capabilities Index, we compare simple models of AI capability trends, finding that the best-fitting model features an acceleration in improvement around April 2024. The rate of frontier improvement nearly doubled, from about 8 points/year before the breakpoint to 15 points/year after. HCAST-based estimates of this inflection point yield similar results.
Data
We use data from the Epoch Capabilities Index (ECI), which itself builds off the paper “A Rosetta Stone for AI Benchmarks”. The current implementation of ECI filters out all models released before 2023, and further filters any models with less than four ECI-included benchmark scores. These requirements were relaxed for the purpose of improving this retrospective analysis’ scope, allowing earlier models with sparser data to be included.
Specifically, models before 2023 were permitted to have as few as three ECI-included benchmark scores, though all models estimated to be at the frontier had at least 4 benchmarks. This left only two points in the 2019-2020 range, so the data cutoff was extended by only a year to preserve continuity. The analysis spans 149 models over four years, ranging from December 2021 to December 2025.
We found that re-fitting the ECI with extended data changed the models’ rankings very little despite extra data; the largest jump was Claude 3 Haiku (From 97th to 102nd place) and the third-largest jump was a movement by one place. That is, while some granular score shifts may have happened, we are confident that general trends did not change.
Analysis
Upon obtaining the newly fitted ECI scores, we keep a running maximum in order to determine the model with the highest ECI at each timepoint. This yields 17 frontier points.
Next, we fit a two-segment linear model to frontier ECI vs time, with a single breakpoint and a continuity constraint at the breakpoint. Dates are converted to numeric days for regression purposes, and 5000 breakpoint candidates are chosen between the 10th and 90th percentile of dates (assuming there are no breaks near the edges).
For each candidate breakpoint, we fit a single continuous piecewise‑linear model by OLS using a hinge basis. This jointly optimizes both segments under the continuity constraint (the right‑hand intercept is implied, not fit separately). Candidates are scored by residual sum of squares, and the breakpoint with the lowest RSS is selected.
April 8, 2024 was chosen as the best breakpoint candidate. Our fit has an R2 of 0.9653, with a pre-breakpoint slope of 8.337 ECI/yr, and a post-breakpoint slope of 15.459 ECI/yr. The post-breakpoint slope is 1.85x as large, and can be interpreted as a marked change in the rate of AI progress. We compare this breakpoint model to a simple linear OLS regression in the table below:
| AIC | BIC | |
|---|---|---|
| OLS model (k = 2) | 48.8 | 50.4 |
| Two-segment model (k = 4) | 43.6 | 47.0 |
We find that the AIC and BIC are both lower in the two-segment model, implying an Akaike Evidence Ratio of 13.5 and a Bayes factor of 5.7. Across 2000 resampled frontier datasets (same size, sampled with replacement), the two-segment model beat the single-line fit on AIC 90% of the time and on BIC about 80% of the time. This reinforces our belief that the breakpoint signal is fairly robust.
We also obtain 90% confidence intervals via a nonparametric bootstrap, resampling the dataset with replacement and refitting the piecewise model with new frontier points each time.
| 5th percentile estimate | 95th percentile estimate | |
|---|---|---|
| Pre-breakpoint slope | +6 ECI / year | +11 ECI / year |
| Post-breakpoint slope | +13 ECI / year | +18 ECI / year |
| Slope ratio (speedup factor) | 1.3x | 3.1x |
| Breakpoint | 2023-03-17 | 2024-08-19 |
Assumptions and limitations
ECI is a composite index constructed from many underlying benchmarks. While we are confident that it tracks a single underlying factor of capabilities progress, interpreting that factor is a bit difficult. For more information, see our ECI FAQ section.
Several data points in our extended ECI dataset feature very wide confidence intervals for capability estimates. Caution should be used when interpreting point estimates for data before 2023.
While we compare piecewise linear models against a simple linear OLS fit, we do not compare against a model of smooth exponential growth. We leave this for future analysis.
Errata
Previously, we enforced continuity in the segmented model by adjusting the right intercept such that the two lines meet at the breakpoint. However, it was suggested that this method would make goodness of fit worse on the second line—that is, we could find a better overall fit by optimizing both segments simultaneously. Following this feedback, we edited the methodology. This resulted in negligible (~0.1) differences in our estimates of the ECI progress slopes, and the optimal breakpoint shifted from April 9, 2024 to April 8, 2024. METR progress slopes did not change, and an identical breakpoint was chosen by this method.

