A core problem in assessing trends in AI capabilities is that benchmarks tend to saturate within 1-3 years. The Epoch Capabilities Index allows comparisons across a wide range of model capabilities, by stitching together benchmarks.
| Model version ID | Benchmark | Publication date | ECI | Performance | ID | Model | Display name |
|---|---|---|---|---|---|---|---|
| gpt-5-nano-2025-08-07_high | FrontierMath-2025-02-28-Private | 2025-08-07 | 138.81 | 0.08 | gpt-5-nano-2025-08-07_high | GPT-5 nano | GPT-5 nano (high) |
| gpt-5-mini-2025-08-07_medium | OTIS Mock AIME 2024-2025 | 2025-08-07 | 141.96 | 0.78 | gpt-5-mini-2025-08-07_medium | GPT-5 mini | GPT-5 mini (medium) |
| gpt-5-mini-2025-08-07_medium | FrontierMath-2025-02-28-Private | 2025-08-07 | 141.96 | 0.19 | gpt-5-mini-2025-08-07_medium | GPT-5 mini | GPT-5 mini (medium) |
| gemini-2.5-flash-preview-05-20 | OTIS Mock AIME 2024-2025 | 2025-05-20 | 139.37 | 0.71 | gemini-2.5-flash-preview-05-20 | Gemini 2.5 Flash | |
| gemini-2.5-flash-preview-04-17 | OTIS Mock AIME 2024-2025 | 2025-04-17 | 138.57 | 0.73 | gemini-2.5-flash-preview-04-17 | Gemini 2.5 Flash | Gemini 2.5 Flash Preview (Apr 2025) |
| Mistral-7B-v0.1 | GSM8K | 2023-09-27 | 112.02 | 0.54 | Mistral-7B-v0.1 | Mistral 7B | |
| gemma-7b | GSM8K | 2024-02-21 | 111.20 | 0.46 | gemma-7b | Gemma 7B | |
| gpt-5-2025-08-07_medium | OTIS Mock AIME 2024-2025 | 2025-08-07 | 150.00 | 0.87 | gpt-5-2025-08-07_medium | GPT-5 | GPT-5 (medium) |
| gpt-5-2025-08-07_medium | FrontierMath-2025-02-28-Private | 2025-08-07 | 150.00 | 0.25 | gpt-5-2025-08-07_medium | GPT-5 | GPT-5 (medium) |
| Mixtral-8x7B-v0.1 | GSM8K | 2023-12-11 | 118.34 | 0.74 | Mixtral-8x7B-v0.1 | Mixtral 8x7B | |
| Yi-6B | GSM8K | 2023-11-02 | 103.43 | 0.33 | Yi-6B | Yi 6B | |
| falcon-180B | GSM8K | 2023-09-06 | 114.00 | 0.54 | falcon-180B | Falcon-180B | |
| Llama-2-7b | GSM8K | 2023-07-18 | 98.43 | 0.17 | Llama-2-7b | Llama 2-7B | |
| gpt-5-nano-2025-08-07_medium | OTIS Mock AIME 2024-2025 | 2025-08-07 | 139.23 | 0.74 | gpt-5-nano-2025-08-07_medium | GPT-5 nano | GPT-5 nano (medium) |
| gpt-5-nano-2025-08-07_medium | FrontierMath-2025-02-28-Private | 2025-08-07 | 139.23 | 0.07 | gpt-5-nano-2025-08-07_medium | GPT-5 nano | GPT-5 nano (medium) |
| Llama-2-13b | GSM8K | 2023-07-18 | 105.39 | 0.34 | Llama-2-13b | Llama 2-13B | |
| Llama-2-34b | GSM8K | 2023-07-18 | 106.25 | 0.42 | Llama-2-34b | Llama 2-34B | |
| Llama-2-70b-hf | GSM8K | 2023-07-18 | 116.83 | 0.70 | Llama-2-70b-hf | Llama 2-70B | |
| falcon-7b | GSM8K | 2023-04-24 | 95.04 | 0.07 | falcon-7b | Falcon-7B | |
| falcon-40b | GSM8K | 2023-03-15 | 104.41 | 0.25 | falcon-40b | Falcon-40B | |
| mpt-7b | GSM8K | 2023-05-05 | 94.22 | 0.09 | mpt-7b | MPT-7B | |
| chatglm2-6b | GSM8K | 2023-06-24 | 102.11 | 0.32 | chatglm2-6b | ||
| internlm-7b | GSM8K | 2023-07-05 | 101.37 | 0.31 | internlm-7b | ||
| internlm-20b | GSM8K | 2023-09-18 | 113.01 | 0.63 | internlm-20b | ||
| Baichuan-2-7B-Base | GSM8K | 2023-09-20 | 94.28 | 0.25 | Baichuan-2-7B-Base | Baichuan 2-7B | |
| Kimi-K2-Instruct | OTIS Mock AIME 2024-2025 | 2025-07-12 | 134.87 | 0.33 | Kimi-K2-Instruct | Kimi K2 | |
| Kimi-K2-Instruct | FrontierMath-2025-02-28-Private | 2025-07-12 | 134.87 | 0.02 | Kimi-K2-Instruct | Kimi K2 | |
| Baichuan-2-13B-Base | GSM8K | 2023-09-06 | 102.45 | 0.53 | Baichuan-2-13B-Base | Baichuan2-13B | |
| LLaMA-7B | GSM8K | 2023-02-24 | 96.01 | 0.11 | LLaMA-7B | LLaMA-7B | |
| LLaMA-13B | GSM8K | 2023-02-27 | 100.56 | 0.21 | LLaMA-13B | LLaMA-13B | |
| LLaMA-33B | GSM8K | 2023-02-27 | 108.68 | 0.44 | LLaMA-33B | LLaMA-33B | |
| LLaMA-65B | GSM8K | 2023-02-24 | 111.71 | 0.54 | LLaMA-65B | LLaMA-65B | |
| StableBeluga2 | GSM8K | 2023-07-20 | 117.83 | 0.70 | StableBeluga2 | Stable Beluga 2 | |
| Qwen-1_8B | GSM8K | 2023-11-30 | 92.39 | 0.21 | Qwen-1_8B | Qwen-1.8B | |
| Qwen-7B | GSM8K | 2023-09-28 | 106.94 | 0.52 | Qwen-7B | Qwen-7B | |
| Qwen-14B | GSM8K | 2023-09-28 | 113.41 | 0.61 | Qwen-14B | Qwen-14B | |
| Nemotron-4 15B | GSM8K | 2024-02-26 | 106.01 | 0.46 | Nemotron-4 15B | Nemotron-4 15B | |
| claude-opus-4-1-20250805_27K | OTIS Mock AIME 2024-2025 | 2025-08-05 | 140.63 | 0.69 | claude-opus-4-1-20250805_27K | Claude Opus 4.1 | |
| claude-opus-4-1-20250805_27K | FrontierMath-2025-02-28-Private | 2025-08-05 | 140.63 | 0.07 | claude-opus-4-1-20250805_27K | Claude Opus 4.1 | |
| mpt-30b | GSM8K | 2023-06-22 | 99.10 | 0.16 | mpt-30b | MPT-30B | |
| gemma-2b | GSM8K | 2024-02-21 | 93.92 | 0.18 | gemma-2b | Gemma 2B | |
| Qwen2.5-Coder-0.5B | GSM8K | 2024-09-18 | 84.51 | 0.35 | Qwen2.5-Coder-0.5B | ||
| Qwen2.5-Coder-1.5B | GSM8K | 2024-09-18 | 104.16 | 0.66 | Qwen2.5-Coder-1.5B | Qwen2.5-Coder (1.5B) | |
| Qwen2.5-Coder-3B | GSM8K | 2024-09-18 | 109.07 | 0.76 | Qwen2.5-Coder-3B | ||
| Qwen2.5-Coder-7B | GSM8K | 2024-09-18 | 113.60 | 0.84 | Qwen2.5-Coder-7B | Qwen2.5-Coder (7B) | |
| claude-opus-4-1-20250805 | OTIS Mock AIME 2024-2025 | 2025-08-05 | 138.59 | 0.40 | claude-opus-4-1-20250805 | Claude Opus 4.1 | |
| claude-opus-4-1-20250805 | FrontierMath-2025-02-28-Private | 2025-08-05 | 138.59 | 0.06 | claude-opus-4-1-20250805 | Claude Opus 4.1 | |
| Qwen2.5-Coder-14B | GSM8K | 2024-09-18 | 117.00 | 0.89 | Qwen2.5-Coder-14B | ||
| Qwen2.5-Coder-32B | GSM8K | 2024-09-18 | 119.84 | 0.91 | Qwen2.5-Coder-32B | Qwen2.5-Coder (32B) | |
| vicuna-13b-v1.1 | GSM8K | 2023-04-12 | 99.94 | 0.28 | vicuna-13b-v1.1 | ||
| gemini-2.5-pro-preview-06-05 | FrontierMath-2025-02-28-Private | 2025-06-05 | 146.24 | 0.10 | gemini-2.5-pro-preview-06-05 | Gemini 2.5 Pro | Gemini 2.5 Pro Preview (Jun 2025) |
| Baichuan-7B | GSM8K | 2023-06-01 | 92.42 | 0.09 | Baichuan-7B | Baichuan1-7B | |
| DeepSeek-R1-0528 | OTIS Mock AIME 2024-2025 | 2025-05-28 | 139.70 | 0.66 | DeepSeek-R1-0528 | DeepSeek-R1 | DeepSeek-R1 (May 2025) |
| gpt-5-mini-2025-08-07_high | OTIS Mock AIME 2024-2025 | 2025-08-07 | 144.11 | 0.87 | gpt-5-mini-2025-08-07_high | GPT-5 mini | GPT-5 mini (high) |
| gpt-5-mini-2025-08-07_high | FrontierMath-2025-02-28-Private | 2025-08-07 | 144.11 | 0.20 | gpt-5-mini-2025-08-07_high | GPT-5 mini | GPT-5 mini (high) |
| claude-3-7-sonnet-20250219_64K | OTIS Mock AIME 2024-2025 | 2025-02-24 | 139.22 | 0.58 | claude-3-7-sonnet-20250219_64K | Claude 3.7 Sonnet | Claude 3.7 Sonnet (64k thinking) |
| claude-3-7-sonnet-20250219_64K | FrontierMath-2025-02-28-Private | 2025-02-24 | 139.22 | 0.03 | claude-3-7-sonnet-20250219_64K | Claude 3.7 Sonnet | Claude 3.7 Sonnet (64k thinking) |
| grok-3-mini-beta_high | OTIS Mock AIME 2024-2025 | 2025-04-09 | 140.05 | 0.78 | grok-3-mini-beta_high | Grok-3 mini | |
| grok-3-mini-beta_high | FrontierMath-2025-02-28-Private | 2025-04-09 | 140.05 | 0.06 | grok-3-mini-beta_high | Grok-3 mini | |
| DeepSeek-R1 | OTIS Mock AIME 2024-2025 | 2025-01-20 | 138.04 | 0.53 | DeepSeek-R1 | DeepSeek-R1 | |
| grok-3-beta | OTIS Mock AIME 2024-2025 | 2025-04-09 | 138.07 | 0.56 | grok-3-beta | Grok 3 | |
| grok-3-beta | FrontierMath-2025-02-28-Private | 2025-04-09 | 138.07 | 0.04 | grok-3-beta | Grok 3 | |
| claude-sonnet-4-20250514_32K | OTIS Mock AIME 2024-2025 | 2025-05-22 | 140.19 | 0.71 | claude-sonnet-4-20250514_32K | Claude Sonnet 4 | |
| claude-opus-4-20250514_16K | OTIS Mock AIME 2024-2025 | 2025-05-22 | 139.94 | 0.60 | claude-opus-4-20250514_16K | Claude Opus 4 | |
| claude-sonnet-4-20250514 | OTIS Mock AIME 2024-2025 | 2025-05-22 | 136.30 | 0.29 | claude-sonnet-4-20250514 | Claude Sonnet 4 | |
| claude-sonnet-4-20250514 | FrontierMath-2025-02-28-Private | 2025-05-22 | 136.30 | 0.04 | claude-sonnet-4-20250514 | Claude Sonnet 4 | |
| claude-opus-4-20250514 | OTIS Mock AIME 2024-2025 | 2025-05-22 | 138.36 | 0.42 | claude-opus-4-20250514 | Claude Opus 4 | |
| claude-opus-4-20250514 | FrontierMath-2025-02-28-Private | 2025-05-22 | 138.36 | 0.05 | claude-opus-4-20250514 | Claude Opus 4 | |
| mistral-medium-2505 | OTIS Mock AIME 2024-2025 | 2025-05-07 | 134.27 | 0.32 | mistral-medium-2505 | Mistral Medium 3 | |
| mistral-medium-2505 | FrontierMath-2025-02-28-Private | 2025-05-07 | 134.27 | 0.00 | mistral-medium-2505 | Mistral Medium 3 | |
| o4-mini-2025-04-16_high | OTIS Mock AIME 2024-2025 | 2025-04-16 | 143.53 | 0.82 | o4-mini-2025-04-16_high | o4-mini | o4-mini (high) |
| o4-mini-2025-04-16_high | FrontierMath-2025-02-28-Private | 2025-04-16 | 143.53 | 0.17 | o4-mini-2025-04-16_high | o4-mini | o4-mini (high) |
| gpt-5-2025-08-07_high | OTIS Mock AIME 2024-2025 | 2025-08-07 | 149.63 | 0.91 | gpt-5-2025-08-07_high | GPT-5 | GPT-5 (high) |
| gpt-5-2025-08-07_high | FrontierMath-2025-02-28-Private | 2025-08-07 | 149.63 | 0.27 | gpt-5-2025-08-07_high | GPT-5 | GPT-5 (high) |
| o3-2025-04-16_high | OTIS Mock AIME 2024-2025 | 2025-04-16 | 144.25 | 0.84 | o3-2025-04-16_high | o3 | o3 (high) |
| o3-2025-04-16_high | FrontierMath-2025-02-28-Private | 2025-04-16 | 144.25 | 0.10 | o3-2025-04-16_high | o3 | o3 (high) |
| gpt-4.1-mini-2025-04-14 | OTIS Mock AIME 2024-2025 | 2025-04-14 | 134.65 | 0.45 | gpt-4.1-mini-2025-04-14 | GPT-4.1 mini | GPT-4.1 mini |
| gpt-4.1-mini-2025-04-14 | FrontierMath-2025-02-28-Private | 2025-04-14 | 134.65 | 0.05 | gpt-4.1-mini-2025-04-14 | GPT-4.1 mini | GPT-4.1 mini |
| gpt-4.1-nano-2025-04-14 | OTIS Mock AIME 2024-2025 | 2025-04-14 | 129.87 | 0.29 | gpt-4.1-nano-2025-04-14 | GPT-4.1 nano | GPT-4.1 nano |
| gpt-4.1-nano-2025-04-14 | FrontierMath-2025-02-28-Private | 2025-04-14 | 129.87 | 0.01 | gpt-4.1-nano-2025-04-14 | GPT-4.1 nano | GPT-4.1 nano |
| gpt-4.1-2025-04-14 | OTIS Mock AIME 2024-2025 | 2025-04-14 | 136.25 | 0.38 | gpt-4.1-2025-04-14 | GPT-4.1 | GPT-4.1 |
| gpt-4.1-2025-04-14 | FrontierMath-2025-02-28-Private | 2025-04-14 | 136.25 | 0.06 | gpt-4.1-2025-04-14 | GPT-4.1 | GPT-4.1 |
| grok-3-mini-beta_low | OTIS Mock AIME 2024-2025 | 2025-04-09 | 137.50 | 0.62 | grok-3-mini-beta_low | Grok-3 mini | |
| grok-3-mini-beta_low | FrontierMath-2025-02-28-Private | 2025-04-09 | 137.50 | 0.03 | grok-3-mini-beta_low | Grok-3 mini | |
| Llama-4-Scout-17B-16E-Instruct | OTIS Mock AIME 2024-2025 | 2025-04-05 | 128.88 | 0.08 | Llama-4-Scout-17B-16E-Instruct | Llama 4 Scout | |
| Llama-4-Scout-17B-16E-Instruct | FrontierMath-2025-02-28-Private | 2025-04-05 | 128.88 | 0.00 | Llama-4-Scout-17B-16E-Instruct | Llama 4 Scout | |
| Llama-4-Maverick-17B-128E-Instruct-FP8 | OTIS Mock AIME 2024-2025 | 2025-04-05 | 133.56 | 0.21 | Llama-4-Maverick-17B-128E-Instruct-FP8 | Llama 4 Maverick | Llama 4 Maverick (FP8) |
| Llama-4-Maverick-17B-128E-Instruct-FP8 | FrontierMath-2025-02-28-Private | 2025-04-05 | 133.56 | 0.01 | Llama-4-Maverick-17B-128E-Instruct-FP8 | Llama 4 Maverick | Llama 4 Maverick (FP8) |
| qwen-max-2025-01-25 | OTIS Mock AIME 2024-2025 | 2025-01-25 | 132.57 | 0.16 | qwen-max-2025-01-25 | Qwen2.5-Max | |
| qwen-max-2025-01-25 | FrontierMath-2025-02-28-Private | 2025-01-25 | 132.57 | 0.01 | qwen-max-2025-01-25 | Qwen2.5-Max | |
| DeepSeek-V3-0324 | OTIS Mock AIME 2024-2025 | 2025-03-24 | 136.21 | 0.38 | DeepSeek-V3-0324 | DeepSeek-V3 | DeepSeek-V3 (Mar 2025) |
| gpt-4-0314 | GSM8K | 2023-03-14 | 125.69 | 0.92 | gpt-4-0314 | GPT-4 | GPT-4 (Mar 2023) |
| gpt-4-0314 | OTIS Mock AIME 2024-2025 | 2023-03-14 | 125.69 | 0.01 | gpt-4-0314 | GPT-4 | GPT-4 (Mar 2023) |
| mistral-small-2503 | OTIS Mock AIME 2024-2025 | 2025-03-17 | 127.32 | 0.06 | mistral-small-2503 | Mistral Small 3.1 | |
| gemma-3-27b-it | OTIS Mock AIME 2024-2025 | 2025-03-12 | 129.64 | 0.20 | gemma-3-27b-it | Gemma 3 27B | |
| claude-3-5-haiku-20241022 | OTIS Mock AIME 2024-2025 | 2024-10-22 | 127.45 | 0.04 | claude-3-5-haiku-20241022 | Claude 3.5 Haiku | Claude 3.5 Haiku (Oct 2024) |
| claude-3-5-haiku-20241022 | FrontierMath-2025-02-28-Private | 2024-10-22 | 127.45 | 0.00 | claude-3-5-haiku-20241022 | Claude 3.5 Haiku | Claude 3.5 Haiku (Oct 2024) |
| claude-3-7-sonnet-20250219_32K | OTIS Mock AIME 2024-2025 | 2025-02-24 | 139.08 | 0.53 | claude-3-7-sonnet-20250219_32K | Claude 3.7 Sonnet | Claude 3.7 Sonnet (32k thinking) |
| claude-3-7-sonnet-20250219_32K | FrontierMath-2025-02-28-Private | 2025-02-24 | 139.08 | 0.03 | claude-3-7-sonnet-20250219_32K | Claude 3.7 Sonnet | Claude 3.7 Sonnet (32k thinking) |
| DeepSeek-R1-Distill-Llama-70B | OTIS Mock AIME 2024-2025 | 2025-01-20 | 136.43 | 0.51 | DeepSeek-R1-Distill-Llama-70B | DeepSeek-R1-Distill-Llama-70B | |
| gpt-4.5-preview-2025-02-27 | OTIS Mock AIME 2024-2025 | 2025-02-27 | 136.49 | 0.38 | gpt-4.5-preview-2025-02-27 | GPT-4.5 | GPT-4.5 Preview (Feb 2025) |
| gpt-4-turbo-2024-04-09 | OTIS Mock AIME 2024-2025 | 2024-04-09 | 127.64 | 0.07 | gpt-4-turbo-2024-04-09 | GPT-4 Turbo | |
| claude-3-7-sonnet-20250219_16K | OTIS Mock AIME 2024-2025 | 2025-02-24 | 138.29 | 0.47 | claude-3-7-sonnet-20250219_16K | Claude 3.7 Sonnet | Claude 3.7 Sonnet (16k thinking) |
| claude-3-7-sonnet-20250219_16K | FrontierMath-2025-02-28-Private | 2025-02-24 | 138.29 | 0.04 | claude-3-7-sonnet-20250219_16K | Claude 3.7 Sonnet | Claude 3.7 Sonnet (16k thinking) |
| mistral-large-2411 | OTIS Mock AIME 2024-2025 | 2024-11-18 | 128.48 | 0.08 | mistral-large-2411 | Mistral Large 2 | |
| mistral-large-2411 | FrontierMath-2025-02-28-Private | 2024-11-18 | 128.48 | 0.00 | mistral-large-2411 | Mistral Large 2 | |
| claude-3-7-sonnet-20250219 | OTIS Mock AIME 2024-2025 | 2025-02-24 | 136.22 | 0.22 | claude-3-7-sonnet-20250219 | Claude 3.7 Sonnet | |
| claude-3-7-sonnet-20250219 | FrontierMath-2025-02-28-Private | 2025-02-24 | 136.22 | 0.03 | claude-3-7-sonnet-20250219 | Claude 3.7 Sonnet | |
| o1-2024-12-17_high | FrontierMath-2025-02-28-Private | 2024-12-17 | 140.65 | 0.09 | o1-2024-12-17_high | o1 | o1 (high) |
| o1-mini-2024-09-12_high | OTIS Mock AIME 2024-2025 | 2024-09-12 | 136.82 | 0.47 | o1-mini-2024-09-12_high | o1-mini | o1-mini (high) |
| o1-mini-2024-09-12_high | FrontierMath-2025-02-28-Private | 2024-09-12 | 136.82 | 0.01 | o1-mini-2024-09-12_high | o1-mini | o1-mini (high) |
| o3-mini-2025-01-31_high | OTIS Mock AIME 2024-2025 | 2025-01-31 | 140.49 | 0.77 | o3-mini-2025-01-31_high | o3-mini | o3-mini (high) |
| o3-mini-2025-01-31_high | FrontierMath-2025-02-28-Private | 2025-01-31 | 140.49 | 0.11 | o3-mini-2025-01-31_high | o3-mini | o3-mini (high) |
| gemini-2.0-flash-thinking-exp-01-21 | OTIS Mock AIME 2024-2025 | 2025-01-21 | 135.44 | 0.58 | gemini-2.0-flash-thinking-exp-01-21 | Gemini 2.0 Flash Thinking | Gemini 2.0 Flash Thinking Exp |
| gemini-2.0-flash-001 | OTIS Mock AIME 2024-2025 | 2025-02-05 | 134.61 | 0.31 | gemini-2.0-flash-001 | Gemini 2.0 Flash | Gemini 2.0 Flash (Feb 2025) |
| gemini-2.0-flash-001 | FrontierMath-2025-02-28-Private | 2025-02-05 | 134.61 | 0.02 | gemini-2.0-flash-001 | Gemini 2.0 Flash | Gemini 2.0 Flash (Feb 2025) |
| gpt-4o-2024-11-20 | OTIS Mock AIME 2024-2025 | 2024-11-20 | 129.08 | 0.06 | gpt-4o-2024-11-20 | GPT-4o | GPT-4o (Nov 2024) |
| gpt-4o-2024-11-20 | FrontierMath-2025-02-28-Private | 2024-11-20 | 129.08 | 0.00 | gpt-4o-2024-11-20 | GPT-4o | GPT-4o (Nov 2024) |
| phi-4 | OTIS Mock AIME 2024-2025 | 2024-12-12 | 130.10 | 0.14 | phi-4 | Phi-4 | |
| o3-mini-2025-01-31_medium | OTIS Mock AIME 2024-2025 | 2025-01-31 | 138.58 | 0.64 | o3-mini-2025-01-31_medium | o3-mini | o3-mini (medium) |
| o3-mini-2025-01-31_medium | FrontierMath-2025-02-28-Private | 2025-01-31 | 138.58 | 0.08 | o3-mini-2025-01-31_medium | o3-mini | o3-mini (medium) |
| claude-sonnet-4-5-20250929_32K | OTIS Mock AIME 2024-2025 | 2025-09-29 | 142.50 | 0.78 | claude-sonnet-4-5-20250929_32K | Claude Sonnet 4.5 | |
| claude-sonnet-4-5-20250929_32K | FrontierMath-2025-02-28-Private | 2025-09-29 | 142.50 | 0.10 | claude-sonnet-4-5-20250929_32K | Claude Sonnet 4.5 | |
| mistral-small-2501 | OTIS Mock AIME 2024-2025 | 2025-01-25 | 126.74 | 0.05 | mistral-small-2501 | Mistral Small 3 | |
| qwen2.5-32b-instruct | OTIS Mock AIME 2024-2025 | 2024-09-17 | 128.44 | 0.07 | qwen2.5-32b-instruct | Qwen2.5-32B | |
| gemini-1.5-flash-001 | GSM8K | 2024-05-23 | 121.98 | 0.82 | gemini-1.5-flash-001 | Gemini 1.5 Flash | Gemini 1.5 Flash (May 2024) |
| gemini-1.5-flash-001 | OTIS Mock AIME 2024-2025 | 2024-05-23 | 121.98 | 0.04 | gemini-1.5-flash-001 | Gemini 1.5 Flash | Gemini 1.5 Flash (May 2024) |
| claude-2.0 | OTIS Mock AIME 2024-2025 | 2023-07-11 | 119.94 | 0.03 | claude-2.0 | Claude 2 | |
| claude-3-opus-20240229 | OTIS Mock AIME 2024-2025 | 2024-02-29 | 126.74 | 0.05 | claude-3-opus-20240229 | Claude 3 Opus | |
| o1-mini-2024-09-12_medium | OTIS Mock AIME 2024-2025 | 2024-09-12 | 135.37 | 0.45 | o1-mini-2024-09-12_medium | o1-mini | o1-mini (medium) |
| o1-mini-2024-09-12_medium | FrontierMath-2025-02-28-Private | 2024-09-12 | 135.37 | 0.02 | o1-mini-2024-09-12_medium | o1-mini | o1-mini (medium) |
| gpt-4o-2024-08-06 | OTIS Mock AIME 2024-2025 | 2024-08-06 | 129.03 | 0.06 | gpt-4o-2024-08-06 | GPT-4o | GPT-4o (Aug 2024) |
| gpt-4o-2024-08-06 | FrontierMath-2025-02-28-Private | 2024-08-06 | 129.03 | 0.00 | gpt-4o-2024-08-06 | GPT-4o | GPT-4o (Aug 2024) |
| o1-preview-2024-09-12 | OTIS Mock AIME 2024-2025 | 2024-09-12 | 134.67 | 0.31 | o1-preview-2024-09-12 | o1-preview | o1-preview |
| claude-3-sonnet-20240229 | OTIS Mock AIME 2024-2025 | 2024-02-29 | 120.07 | 0.03 | claude-3-sonnet-20240229 | Claude 3 Sonnet | |
| claude-3-5-sonnet-20241022 | OTIS Mock AIME 2024-2025 | 2024-10-22 | 133.59 | 0.09 | claude-3-5-sonnet-20241022 | Claude 3.5 Sonnet | Claude 3.5 Sonnet (Oct 2024) |
| claude-3-5-sonnet-20241022 | FrontierMath-2025-02-28-Private | 2024-10-22 | 133.59 | 0.02 | claude-3-5-sonnet-20241022 | Claude 3.5 Sonnet | Claude 3.5 Sonnet (Oct 2024) |
| gemma-2-27b-it | OTIS Mock AIME 2024-2025 | 2024-06-24 | 122.79 | 0.01 | gemma-2-27b-it | Gemma 2 27B | |
| claude-3-haiku-20240307 | OTIS Mock AIME 2024-2025 | 2024-03-07 | 117.85 | 0.02 | claude-3-haiku-20240307 | Claude 3 Haiku | |
| claude-3-5-sonnet-20240620 | OTIS Mock AIME 2024-2025 | 2024-06-20 | 130.00 | 0.07 | claude-3-5-sonnet-20240620 | Claude 3.5 Sonnet | Claude 3.5 Sonnet (Jun 2024) |
| claude-3-5-sonnet-20240620 | FrontierMath-2025-02-28-Private | 2024-06-20 | 130.00 | 0.01 | claude-3-5-sonnet-20240620 | Claude 3.5 Sonnet | Claude 3.5 Sonnet (Jun 2024) |
| gpt-4o-2024-05-13 | OTIS Mock AIME 2024-2025 | 2024-05-13 | 128.38 | 0.06 | gpt-4o-2024-05-13 | GPT-4o | GPT-4o (May 2024) |
| gemini-1.5-flash-002 | OTIS Mock AIME 2024-2025 | 2024-09-24 | 129.76 | 0.16 | gemini-1.5-flash-002 | Gemini 1.5 Flash | |
| gemini-1.5-flash-002 | FrontierMath-2025-02-28-Private | 2024-09-24 | 129.76 | 0.00 | gemini-1.5-flash-002 | Gemini 1.5 Flash | |
| gpt-4-0613 | GSM8K | 2023-06-13 | 122.18 | 0.90 | gpt-4-0613 | GPT-4 | GPT-4 (Jun 2023) |
| gpt-4-0613 | OTIS Mock AIME 2024-2025 | 2023-06-13 | 122.18 | 0.01 | gpt-4-0613 | GPT-4 | GPT-4 (Jun 2023) |
| gemma-2-9b-it | OTIS Mock AIME 2024-2025 | 2024-06-24 | 119.28 | 0.01 | gemma-2-9b-it | Gemma 2 9B | |
| gemini-1.5-pro-001 | OTIS Mock AIME 2024-2025 | 2024-05-24 | 126.39 | 0.07 | gemini-1.5-pro-001 | Gemini 1.5 Pro | |
| gpt-4o-mini-2024-07-18 | GSM8K | 2024-07-18 | 126.07 | 0.91 | gpt-4o-mini-2024-07-18 | GPT-4o mini | |
| gpt-4o-mini-2024-07-18 | OTIS Mock AIME 2024-2025 | 2024-07-18 | 126.07 | 0.07 | gpt-4o-mini-2024-07-18 | GPT-4o mini | |
| qwen3-max-2025-09-23 | OTIS Mock AIME 2024-2025 | 2025-09-24 | 140.77 | 0.73 | qwen3-max-2025-09-23 | Qwen3-Max | Qwen3-Max-Instruct |
| gemini-1.5-pro-002 | OTIS Mock AIME 2024-2025 | 2024-09-24 | 132.33 | 0.23 | gemini-1.5-pro-002 | Gemini 1.5 Pro | |
| Llama-3.1-8B-Instruct | GSM8K | 2024-07-23 | 115.59 | 0.82 | Llama-3.1-8B-Instruct | Llama 3.1-8B | |
| Llama-3.1-8B-Instruct | OTIS Mock AIME 2024-2025 | 2024-07-23 | 115.59 | 0.03 | Llama-3.1-8B-Instruct | Llama 3.1-8B | |
| Llama-3.1-70B-Instruct | OTIS Mock AIME 2024-2025 | 2024-07-23 | 125.57 | 0.04 | Llama-3.1-70B-Instruct | Llama 3.1-70B | |
| Llama-3.1-405B-Instruct | OTIS Mock AIME 2024-2025 | 2024-07-23 | 128.07 | 0.10 | Llama-3.1-405B-Instruct | Llama 3.1-405B | |
| grok-2-1212 | OTIS Mock AIME 2024-2025 | 2024-12-12 | 130.29 | 0.12 | grok-2-1212 | Grok-2 | |
| grok-2-1212 | FrontierMath-2025-02-28-Private | 2024-12-12 | 130.29 | 0.01 | grok-2-1212 | Grok-2 | |
| Meta-Llama-3-70B-Instruct | OTIS Mock AIME 2024-2025 | 2024-04-18 | 121.84 | 0.04 | Meta-Llama-3-70B-Instruct | Llama 3-70B | |
| Meta-Llama-3-8B-Instruct | OTIS Mock AIME 2024-2025 | 2024-04-18 | 116.40 | 0.01 | Meta-Llama-3-8B-Instruct | Llama 3-8B | |
| qwen2.5-72b-instruct | OTIS Mock AIME 2024-2025 | 2024-09-19 | 129.69 | 0.08 | qwen2.5-72b-instruct | Qwen2.5-72B | |
| claude-sonnet-4-5-20250929 | OTIS Mock AIME 2024-2025 | 2025-09-29 | 138.07 | 0.36 | claude-sonnet-4-5-20250929 | Claude Sonnet 4.5 | |
| claude-sonnet-4-5-20250929 | FrontierMath-2025-02-28-Private | 2025-09-29 | 138.07 | 0.05 | claude-sonnet-4-5-20250929 | Claude Sonnet 4.5 | |
| gemini-1.0-pro-001 | OTIS Mock AIME 2024-2025 | 2024-02-15 | 117.23 | 0.01 | gemini-1.0-pro-001 | Gemini 1.0 Pro | |
| mistral-large-2407 | OTIS Mock AIME 2024-2025 | 2024-07-24 | 127.36 | 0.09 | mistral-large-2407 | Mistral Large 2 | |
| Llama-3.2-90B-Vision-Instruct | OTIS Mock AIME 2024-2025 | 2024-09-24 | 125.44 | 0.03 | Llama-3.2-90B-Vision-Instruct | Llama 3.2 90B | |
| Llama-3.3-70B-Instruct | OTIS Mock AIME 2024-2025 | 2024-12-06 | 127.46 | 0.05 | Llama-3.3-70B-Instruct | Llama 3.3 70B | |
| o1-2024-12-17_medium | OTIS Mock AIME 2024-2025 | 2024-12-17 | 140.83 | 0.73 | o1-2024-12-17_medium | o1 | o1 (medium) |
| DeepSeek-V3 | OTIS Mock AIME 2024-2025 | 2024-12-26 | 132.39 | 0.16 | DeepSeek-V3 | DeepSeek-V3 | |
| DeepSeek-V3 | FrontierMath-2025-02-28-Private | 2024-12-26 | 132.39 | 0.02 | DeepSeek-V3 | DeepSeek-V3 | |
| grok-4-0709 | OTIS Mock AIME 2024-2025 | 2025-07-09 | 145.90 | 0.84 | grok-4-0709 | Grok 4 | |
| grok-4-0709 | FrontierMath-2025-02-28-Private | 2025-07-09 | 145.90 | 0.12 | grok-4-0709 | Grok 4 | |
| o4-mini-2025-04-16_medium | FrontierMath-2025-02-28-Private | 2025-04-16 | 141.89 | 0.19 | o4-mini-2025-04-16_medium | o4-mini | o4-mini (medium) |
| o3-2025-04-16_medium | FrontierMath-2025-02-28-Private | 2025-04-16 | 144.14 | 0.10 | o3-2025-04-16_medium | o3 | o3 (medium) |
The ECI uses an abstract scale, but scores can be interpreted by calculating the expected performance on individual benchmarks. In the chart above, we show expected benchmark performance across a range of ECI values for GSM8K, OTIS Mock AIME 2024-2025, and FrontierMath Tier 1-3.
Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.
Learn more about this graph
One way of interpreting an Epoch Capabilities Index (ECI) score is by calculating the expected performance on individual benchmarks based only on the ECI score. We plot this relationship for GSM8K, OTIS Mock AIME 2024-2025, and FrontierMath Tier 1-3, three benchmarks that vary considerably in difficulty. We also show 90% confidence intervals as a shaded region.
Data
The ECI is calculated using data from Epoch’s Benchmarking Hub. We utilize both internally run evaluations, as well as benchmark creator- and model developer-reported evaluations. Currently, the ECI uses 1123 distinct evaluations, covering 147 models and 39 underlying benchmarks. More details on the methodology used for the ECI can be found here.
Analysis
ECI is based on a statistical model that estimates three types of parameters. Each AI model is given an estimated capability (which we use to produce the ECI), and each benchmark is given location and slope parameters, indicating the model’s overall difficulty and the range of difficulties among its constituent problems, respectively.
These parameters are estimated by fitting the observed data to the following formula:
$$ \textrm{performance}(m,b) = \sigma(\alpha_b [C_m - D_b]) $$
Where \(C_m\) is capability for model \(m\) (i.e. ECI), \(D_b\) is overall difficulty for benchmark \(b\), and \(\alpha_b\) is the “slope” for benchmark \(b\).
To back out the expected performance on each benchmark across a range of ECI scores, we simply plug in the estimated \(\alpha_b\) and \(D_b\) values for each benchmark, and plot how performance changes as \(C_m\) varies.
We calculate 90% confidence intervals by bootstrap resampling from our observed scores, fitting a new model for each resample, and calculating the 5th and 95th percentile for expected scores across samples.
Limitations
This analysis focuses on the expected score on several benchmarks for a given ECI score. The plotted 90% confidence intervals represent a confidence interval over the expected performance, not a prediction interval for the performance obtainable for a given ECI. In practice, individual models may perform above or below the 90% confidence interval.
More information on the technical details and limitations of ECI can be found on the ECI tab of the AI Benchmarking hub.

