Across benchmarks measuring skills in research-level math, agentic coding, visual understanding, common sense reasoning, and more, AI capabilities have grown rapidly and consistently over the last 12 months.
| Benchmark | Model | Organization | Version release date | Score |
|---|---|---|---|---|
| GPQA Diamond | Grok 4 | xAI | 2025-07-09 | 0.87 |
| GPQA Diamond | Gemini 1.5 Flash | Google DeepMind | 2024-05-23 | 0.40 |
| GPQA Diamond | Claude 2 | Anthropic | 2023-07-11 | 0.35 |
| GPQA Diamond | Claude 3 Opus | Anthropic | 2024-02-29 | 0.47 |
| GPQA Diamond | o1-mini | OpenAI | 2024-09-12 | 0.60 |
| GPQA Diamond | GPT-4o | OpenAI | 2024-08-06 | 0.49 |
| GPQA Diamond | o1-preview | OpenAI | 2024-09-12 | 0.50 |
| GPQA Diamond | Claude 3 Sonnet | Anthropic | 2024-02-29 | 0.41 |
| GPQA Diamond | Claude 3.5 Sonnet | Anthropic | 2024-10-22 | 0.55 |
| GPQA Diamond | GPT-4 Turbo | OpenAI | 2024-01-25 | 0.42 |
| GPQA Diamond | Claude 2.1 | Anthropic | 2023-11-21 | 0.33 |
| GPQA Diamond | Gemma 2 27B | Google DeepMind | 2024-06-24 | 0.36 |
| GPQA Diamond | Claude 3 Haiku | Anthropic | 2024-03-07 | 0.36 |
| GPQA Diamond | Claude 3.5 Sonnet | Anthropic | 2024-06-20 | 0.54 |
| GPQA Diamond | GPT-4o | OpenAI | 2024-05-13 | 0.49 |
| GPQA Diamond | Gemini 1.5 Flash | Google DeepMind | 2024-09-24 | 0.47 |
| GPQA Diamond | GPT-4 | OpenAI | 2023-06-13 | 0.31 |
| GPQA Diamond | Gemma 2 9B | Google DeepMind | 2024-06-24 | 0.27 |
| GPQA Diamond | GPT-3.5 Turbo | OpenAI | 2023-11-06 | 0.28 |
| GPQA Diamond | Gemini 1.5 Pro | Google DeepMind | 2024-05-24 | 0.46 |
| GPQA Diamond | GPT-4 Turbo | OpenAI | 2023-11-06 | 0.42 |
| GPQA Diamond | GPT-4o mini | OpenAI | 2024-07-18 | 0.38 |
| GPQA Diamond | Gemini 1.5 Pro | Google DeepMind | 2024-09-24 | 0.57 |
| GPQA Diamond | GPT-3.5 Turbo | OpenAI | 2024-01-25 | 0.27 |
| GPQA Diamond | Llama 3.1-8B | Meta AI | 2024-07-23 | 0.26 |
| GPQA Diamond | Llama 3.1-70B | Meta AI | 2024-07-23 | 0.44 |
| GPQA Diamond | Llama 3.1-405B | Meta AI | 2024-07-23 | 0.51 |
| GPQA Diamond | Yi-1.5-34B | 01.AI | 2024-05-13 | 0.32 |
| GPQA Diamond | Yi-34B | 01.AI | 2023-11-22 | 0.15 |
| GPQA Diamond | Grok-2 | xAI | 2024-12-12 | 0.54 |
| GPQA Diamond | Qwen2-72B | Alibaba | 2024-06-07 | 0.41 |
| GPQA Diamond | Qwen1.5-72B | Alibaba | 2024-02-04 | 0.29 |
| GPQA Diamond | Qwen1.5-32B | Alibaba | 2024-04-03 | 0.31 |
| GPQA Diamond | Hermes 2 Theta Llama-3 70B | Nous Research,Arcee AI | 2024-06-20 | 0.37 |
| GPQA Diamond | Llama 2-70B | Meta AI | 2023-07-18 | 0.26 |
| GPQA Diamond | Llama 3-70B | Meta AI | 2024-04-18 | 0.41 |
| GPQA Diamond | Llama 3-8B | Meta AI | 2024-04-18 | 0.26 |
| GPQA Diamond | Mistral 7B | Mistral AI | 2024-05-27 | 0.15 |
| GPQA Diamond | DeepSeek LLM 67B | DeepSeek | 2023-11-29 | 0.25 |
| GPQA Diamond | Mixtral 8x7B | Mistral AI | 2023-12-11 | 0.31 |
| GPQA Diamond | Qwen2.5-72B | Alibaba | 2024-09-19 | 0.49 |
| GPQA Diamond | WizardLM-2 8x22B | Microsoft | 2024-04-15 | 0.43 |
| GPQA Diamond | DBRX | Databricks | 2024-03-27 | 0.33 |
| GPQA Diamond | Gemini 1.0 Pro | Google DeepMind | 2024-02-15 | 0.34 |
| GPQA Diamond | Ministral 3B | Mistral AI | 2024-10-16 | 0.25 |
| GPQA Diamond | Mistral Large | Mistral AI | 2024-02-26 | 0.39 |
| GPQA Diamond | Mixtral 8x22B | Mistral AI | 2024-04-17 | 0.34 |
| GPQA Diamond | Ministral 8B | Mistral AI | 2024-10-16 | 0.27 |
| GPQA Diamond | Mistral Large 2 | Mistral AI | 2024-07-24 | 0.49 |
| GPQA Diamond | Mixtral 8x7B | Mistral AI | 2023-12-11 | 0.30 |
| GPQA Diamond | Mistral 7B | Mistral AI | 2023-09-27 | 0.13 |
| GPQA Diamond | Mistral NeMo | Mistral AI | 2024-07-18 | 0.30 |
| GPQA Diamond | Llama 3.2 90B | Meta AI | 2024-09-24 | 0.41 |
| GPQA Diamond | Tulu 3 (Tülu 3) 70B | Allen Institute for AI,University of Washington | 2024-11-21 | 0.46 |
| GPQA Diamond | Eurus-2-7B-PRIME | Tsinghua University,University of Illinois Urbana-Champaign (UIUC),Shanghai AI Lab,Peking University,Shanghai Jiao Tong University,CUHK Shenzhen Research Institute | 2024-12-31 | 0.34 |
| GPQA Diamond | Llama 3.3 70B | Meta AI | 2024-12-06 | 0.47 |
| GPQA Diamond | o1 | OpenAI | 2024-12-17 | 0.76 |
| GPQA Diamond | DeepSeek-V3 | DeepSeek | 2024-12-26 | 0.57 |
| GPQA Diamond | Qwen2.5-32B | Alibaba | 2024-09-17 | 0.46 |
| GPQA Diamond | Mistral Small 3 | Mistral AI | 2025-01-25 | 0.45 |
| GPQA Diamond | o3-mini | OpenAI | 2025-01-31 | 0.74 |
| GPQA Diamond | phi-3-medium 14B | Microsoft | 2024-04-23 | 0.28 |
| GPQA Diamond | Phi-4 | Microsoft Research | 2024-12-12 | 0.56 |
| GPQA Diamond | GPT-4o | OpenAI | 2024-11-20 | 0.48 |
| GPQA Diamond | Gemini 2.0 Flash | Google DeepMind,Google | 2025-02-05 | 0.64 |
| GPQA Diamond | Gemini 2.0 Flash Thinking | Google DeepMind,Google | 2025-01-21 | 0.57 |
| GPQA Diamond | Gemini 2.0 Pro | Google DeepMind | 2025-02-05 | 0.66 |
| GPQA Diamond | o3-mini | OpenAI | 2025-01-31 | 0.77 |
| GPQA Diamond | o1-mini | OpenAI | 2024-09-12 | 0.62 |
| GPQA Diamond | o1 | OpenAI | 2024-12-17 | 0.77 |
| GPQA Diamond | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 0.66 |
| GPQA Diamond | Mistral Large 2 | Mistral AI | 2024-11-18 | 0.51 |
| GPQA Diamond | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 0.77 |
| GPQA Diamond | GPT-4 Turbo | OpenAI | 2024-04-09 | 0.47 |
| GPQA Diamond | GPT-4.5 | OpenAI | 2025-02-27 | 0.69 |
| GPQA Diamond | DeepSeek-R1-Distill-Llama-70B | DeepSeek | 2025-01-20 | 0.56 |
| GPQA Diamond | DeepSeek-R1-Distill-Qwen-14B | DeepSeek | 2025-01-20 | 0.45 |
| GPQA Diamond | Gemini 1.5 Flash 8B | Google DeepMind | 2024-10-03 | 0.33 |
| GPQA Diamond | Claude 3.5 Haiku | Anthropic | 2024-10-22 | 0.38 |
| GPQA Diamond | Gemma 3 27B | Google DeepMind | 2025-03-12 | 0.49 |
| GPQA Diamond | Mistral Small 3.1 | Mistral AI | 2025-03-17 | 0.47 |
| GPQA Diamond | Gemini 2.5 Pro | Google DeepMind | 2025-03-25 | 0.84 |
| GPQA Diamond | DeepSeek-V3 | DeepSeek | 2025-03-24 | 0.68 |
| GPQA Diamond | Qwen2.5-Max | Alibaba | 2025-01-25 | 0.56 |
| GPQA Diamond | Qwen Plus | Alibaba | 2025-01-25 | 0.48 |
| GPQA Diamond | Qwen-Turbo | Alibaba | 2024-11-01 | 0.42 |
| GPQA Diamond | Llama 4 Maverick | Meta AI | 2025-04-05 | 0.67 |
| GPQA Diamond | Llama 4 Scout | Meta AI | 2025-04-05 | 0.52 |
| GPQA Diamond | Grok-3 mini | xAI | 2025-04-09 | 0.76 |
| GPQA Diamond | QWQ-Plus | Alibaba | 2025-04-08 | 0.65 |
| GPQA Diamond | GPT-4.1 | OpenAI | 2025-04-14 | 0.67 |
| GPQA Diamond | GPT-4.1 nano | OpenAI | 2025-04-14 | 0.49 |
| GPQA Diamond | GPT-4.1 mini | OpenAI | 2025-04-14 | 0.66 |
| GPQA Diamond | o3 | OpenAI | 2025-04-16 | 0.82 |
| GPQA Diamond | o4-mini | OpenAI | 2025-04-16 | 0.80 |
| GPQA Diamond | Mistral Medium 3 | Mistral AI | 2025-05-07 | 0.60 |
| GPQA Diamond | Claude Opus 4 | Anthropic | 2025-05-22 | 0.69 |
| GPQA Diamond | Claude Sonnet 4 | Anthropic | 2025-05-22 | 0.67 |
| GPQA Diamond | Claude Sonnet 4 | Anthropic | 2025-05-22 | 0.76 |
| GPQA Diamond | Claude Opus 4 | Anthropic | 2025-05-22 | 0.76 |
| GPQA Diamond | Claude Sonnet 4 | Anthropic | 2025-05-22 | 0.78 |
| GPQA Diamond | Grok 3 | xAI | 2025-04-09 | 0.76 |
| GPQA Diamond | Claude Sonnet 4 | Anthropic | 2025-05-22 | 0.79 |
| GPQA Diamond | DeepSeek-R1 | DeepSeek | 2025-01-20 | 0.72 |
| GPQA Diamond | Grok-3 mini | xAI | 2025-04-09 | 0.76 |
| GPQA Diamond | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 0.80 |
| GPQA Diamond | DeepSeek-R1 | DeepSeek | 2025-05-28 | 0.76 |
| GPQA Diamond | Gemini 2.5 Pro | Google DeepMind | 2025-05-06 | 0.67 |
| GPQA Diamond | Qwen3-235B-A22B | Alibaba | 2025-04-29 | 0.71 |
| GPQA Diamond | Gemini 2.5 Pro | Google DeepMind | 2025-06-05 | 0.85 |
| GPQA Diamond | Magistral Small 1.1 | Mistral AI | 2025-06-10 | 0.48 |
| GPQA Diamond | Claude Opus 4.1 | Anthropic | 2025-08-05 | 0.73 |
| GPQA Diamond | Claude Opus 4.1 | Anthropic | 2025-08-05 | 0.77 |
| GPQA Diamond | Claude Opus 4.1 | Anthropic | 2025-08-05 | 0.77 |
| GPQA Diamond | Kimi K2 | Moonshot | 2025-07-12 | 0.55 |
| GPQA Diamond | GPT-5 mini | OpenAI | 2025-08-07 | 0.74 |
| GPQA Diamond | GPT-5 nano | OpenAI | 2025-08-07 | 0.66 |
| GPQA Diamond | GPT-5 nano | OpenAI | 2025-08-07 | 0.67 |
| GPQA Diamond | GPT-5 | OpenAI | 2025-08-07 | 0.85 |
| GPQA Diamond | GPT-5 mini | OpenAI | 2025-08-07 | 0.72 |
| GPQA Diamond | GPT-5 | OpenAI | 2025-08-07 | 0.85 |
| GPQA Diamond | Claude Sonnet 4.5 | Anthropic | 2025-09-29 | 0.74 |
| Aider Polyglot | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 64.90 |
| Aider Polyglot | o1 | OpenAI | 2024-12-17 | 61.70 |
| Aider Polyglot | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 60.40 |
| Aider Polyglot | o3-mini | OpenAI | 2025-01-31 | 60.40 |
| Aider Polyglot | DeepSeek-R1 | DeepSeek | 2025-01-20 | 56.90 |
| Aider Polyglot | o3-mini | OpenAI | 2025-01-31 | 53.80 |
| Aider Polyglot | Claude 3.5 Sonnet | Anthropic | 2024-10-22 | 51.60 |
| Aider Polyglot | DeepSeek-V3 | DeepSeek | 2024-12-26 | 48.40 |
| Aider Polyglot | GPT-4.5 | OpenAI | 2025-02-27 | 44.90 |
| Aider Polyglot | Gemini 2.0 Flash | Google DeepMind,Google | 2024-12-06 | 38.20 |
| Aider Polyglot | Gemini 2.0 Pro | Google DeepMind | 2025-02-05 | 35.60 |
| Aider Polyglot | o1-mini | OpenAI | 2024-09-12 | 32.90 |
| Aider Polyglot | Claude 3.5 Haiku | Anthropic | 2024-10-22 | 28.00 |
| Aider Polyglot | GPT-4o | OpenAI | 2025-01-29 | 27.10 |
| Aider Polyglot | GPT-4o | OpenAI | 2024-08-06 | 23.10 |
| Aider Polyglot | Gemini 2.0 Flash | Google DeepMind,Google | 2024-12-11 | 22.20 |
| Aider Polyglot | Qwen2.5-Max | Alibaba | 2025-01-25 | 21.80 |
| Aider Polyglot | QwQ-32B | Alibaba | 2025-03-05 | 20.90 |
| Aider Polyglot | Gemini 2.0 Flash Thinking | Google DeepMind,Google | 2025-01-21 | 18.20 |
| Aider Polyglot | GPT-4o | OpenAI | 2024-11-20 | 18.20 |
| Aider Polyglot | DeepSeek-V2.5 | DeepSeek | 2024-09-06 | 17.80 |
| Aider Polyglot | Qwen2.5-Coder (32B) | Alibaba | 2024-11-21 | 16.40 |
| Aider Polyglot | Yi-Lightning | 01.AI | 2024-12-02 | 12.90 |
| Aider Polyglot | Cohere Command A | Cohere | 2025-03-13 | 12.00 |
| Aider Polyglot | Codestral | Mistral AI | 2025-01-13 | 11.10 |
| Aider Polyglot | Qwen2.5-Coder (32B) | Alibaba | 2024-11-21 | 8.00 |
| Aider Polyglot | Gemma 3 27B | Google DeepMind | 2025-03-12 | 4.90 |
| Aider Polyglot | GPT-4o mini | OpenAI | 2024-07-18 | 3.60 |
| Aider Polyglot | o3 | OpenAI | 2025-04-16 | 79.60 |
| Aider Polyglot | o4-mini | OpenAI | 2025-04-16 | 72.00 |
| Aider Polyglot | Gemini 2.5 Pro | Google DeepMind | 2025-03-25 | 72.90 |
| Aider Polyglot | DeepSeek-R1 | DeepSeek | 2025-05-28 | 71.40 |
| Aider Polyglot | DeepSeek-V3 | DeepSeek | 2025-03-24 | 55.10 |
| Aider Polyglot | Grok 3 | xAI | 2025-04-09 | 53.30 |
| Aider Polyglot | GPT-4.1 | OpenAI | 2025-04-14 | 52.40 |
| Aider Polyglot | Grok-3 mini | xAI | 2025-04-09 | 49.30 |
| Aider Polyglot | Gemini 2.5 Flash | Google DeepMind | 2025-04-17 | 47.10 |
| Aider Polyglot | GPT-4o | OpenAI | 2025-03-27 | 45.30 |
| Aider Polyglot | GPT-4.1 nano | OpenAI | 2025-04-14 | 8.90 |
| Aider Polyglot | GPT-4.1 mini | OpenAI | 2025-04-14 | 32.40 |
| Aider Polyglot | Llama 4 Maverick | Meta AI | 2025-04-05 | 15.60 |
| Aider Polyglot | Qwen3-235B-A22B | Alibaba | 2025-04-29 | 49.80 |
| Aider Polyglot | Qwen3-32B | Alibaba | 2025-04-29 | 40.00 |
| Aider Polyglot | Claude Opus 4 | Anthropic | 2025-05-22 | 72.00 |
| Aider Polyglot | Claude Opus 4 | Anthropic | 2025-05-22 | 70.70 |
| Aider Polyglot | Claude Sonnet 4 | Anthropic | 2025-05-22 | 61.30 |
| Aider Polyglot | Claude Sonnet 4 | Anthropic | 2025-05-22 | 56.40 |
| Aider Polyglot | Gemini 2.5 Flash | Google DeepMind | 2025-05-20 | 44.00 |
| Aider Polyglot | Gemini 2.5 Flash | Google DeepMind | 2025-05-20 | 55.10 |
| Aider Polyglot | Gemini 2.5 Pro | Google DeepMind | 83.10 | |
| Aider Polyglot | Gemini 2.5 Pro | Google DeepMind | 2025-06-05 | 79.10 |
| Aider Polyglot | Gemini 2.5 Pro | Google DeepMind | 2025-05-06 | 76.90 |
| Aider Polyglot | o3-pro | OpenAI | 2025-06-10 | 84.90 |
| Aider Polyglot | Grok 4 | xAI | 2025-07-09 | 79.60 |
| Aider Polyglot | Kimi K2 | Moonshot | 2025-07-12 | 59.10 |
| Aider Polyglot | Grok-3 mini | xAI | 2025-04-09 | 34.70 |
| Aider Polyglot | o3 | OpenAI | 2025-04-16 | 76.90 |
| VPCT | Claude 3.5 Sonnet | Anthropic | 2024-10-22 | 0.33 |
| VPCT | GPT-4o mini | OpenAI | 2024-07-18 | 0.34 |
| VPCT | o1 | OpenAI | 2024-12-17 | 0.37 |
| VPCT | Gemini 2.5 Flash | Google DeepMind | 2025-04-17 | 0.38 |
| VPCT | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 0.39 |
| VPCT | GPT-4o | OpenAI | 2024-11-20 | 0.40 |
| VPCT | Gemini 2.5 Pro | Google DeepMind | 2025-05-06 | 0.41 |
| VPCT | GPT-4.5 | OpenAI | 2025-02-27 | 0.45 |
| VPCT | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 0.35 |
| VPCT | Gemini 2.5 Pro | Google DeepMind | 2025-04-09 | 0.48 |
| VPCT | o3 | OpenAI | 2025-04-16 | 0.52 |
| VPCT | o4-mini | OpenAI | 2025-04-16 | 0.58 |
| VPCT | Claude Sonnet 4 | Anthropic | 2025-05-22 | 0.30 |
| VPCT | Claude Opus 4 | Anthropic | 2025-05-22 | 0.33 |
| VPCT | Claude Opus 4 | Anthropic | 2025-05-22 | 0.38 |
| VPCT | Gemini 2.5 Pro | Google DeepMind | 2025-06-05 | 0.46 |
| VPCT | Claude Sonnet 4 | Anthropic | 2025-05-22 | 0.34 |
| VPCT | Claude Opus 4.1 | Anthropic | 2025-08-05 | 0.33 |
| VPCT | GPT-5 | OpenAI | 2025-08-07 | 0.66 |
| VPCT | GPT-5 | OpenAI | 2025-08-07 | 0.63 |
| VPCT | GPT-5 mini | OpenAI | 2025-08-07 | 0.40 |
| VPCT | GPT-5 mini | OpenAI | 2025-08-07 | 0.39 |
| VPCT | GPT-5 nano | OpenAI | 2025-08-07 | 0.37 |
| VPCT | GPT-5 nano | OpenAI | 2025-08-07 | 0.35 |
| SimpleBench | Gemini 2.5 Pro | Google DeepMind | 2025-03-25 | 0.52 |
| SimpleBench | o3 | OpenAI | 2025-04-16 | 0.53 |
| SimpleBench | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 0.46 |
| SimpleBench | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 0.45 |
| SimpleBench | o1-preview | OpenAI | 2024-09-12 | 0.42 |
| SimpleBench | Claude 3.5 Sonnet | Anthropic | 2024-10-22 | 0.41 |
| SimpleBench | o1 | OpenAI | 2024-12-17 | 0.40 |
| SimpleBench | o4-mini | OpenAI | 2025-04-16 | 0.39 |
| SimpleBench | o1 | OpenAI | 2024-12-17 | 0.37 |
| SimpleBench | Grok 3 | xAI | 2025-04-09 | 0.36 |
| SimpleBench | GPT-4.5 | OpenAI | 2025-02-27 | 0.35 |
| SimpleBench | Gemini 2.0 Flash | Google DeepMind,Google | 2024-12-06 | 0.31 |
| SimpleBench | Qwen3-235B-A22B | Alibaba | 2025-04-29 | 0.31 |
| SimpleBench | DeepSeek-R1 | DeepSeek | 2025-01-20 | 0.31 |
| SimpleBench | Gemini 2.0 Flash Thinking | Google DeepMind,Google | 2025-01-21 | 0.31 |
| SimpleBench | Llama 4 Maverick | Meta AI | 2025-04-05 | 0.28 |
| SimpleBench | Claude 3.5 Sonnet | Anthropic | 2024-06-20 | 0.28 |
| SimpleBench | DeepSeek-V3 | DeepSeek | 2025-03-24 | 0.27 |
| SimpleBench | Gemini 1.5 Pro | Google DeepMind | 2024-09-24 | 0.27 |
| SimpleBench | GPT-4.1 | OpenAI | 2025-04-14 | 0.27 |
| SimpleBench | GPT-4 Turbo | OpenAI | 2024-04-09 | 0.25 |
| SimpleBench | Claude 3 Opus | Anthropic | 2024-02-29 | 0.24 |
| SimpleBench | Llama 3.1-405B | Meta AI | 2024-07-23 | 0.23 |
| SimpleBench | o3-mini | OpenAI | 2025-01-31 | 0.23 |
| SimpleBench | Grok-2 | xAI | 2024-12-12 | 0.23 |
| SimpleBench | Mistral Large 2 | Mistral AI | 2024-07-24 | 0.23 |
| SimpleBench | Llama 3.3 70B | Meta AI | 2024-12-06 | 0.20 |
| SimpleBench | DeepSeek-V3 | DeepSeek | 2024-12-26 | 0.19 |
| SimpleBench | Gemini 2.0 Flash | Google DeepMind,Google | 2024-12-11 | 0.19 |
| SimpleBench | o1-mini | OpenAI | 2024-09-12 | 0.18 |
| SimpleBench | GPT-4o | OpenAI | 2024-08-06 | 0.18 |
| SimpleBench | Command R+ | Cohere,Cohere for AI | 2024-08-30 | 0.17 |
| SimpleBench | GPT-4o mini | OpenAI | 2024-07-18 | 0.11 |
| SimpleBench | Gemini 2.5 Pro | Google DeepMind | 2025-06-05 | 0.62 |
| SimpleBench | DeepSeek-R1 | DeepSeek | 2025-05-28 | 0.41 |
| SimpleBench | Claude Opus 4 | Anthropic | 2025-05-22 | 0.59 |
| SimpleBench | Claude Sonnet 4 | Anthropic | 2025-05-22 | 0.46 |
| SimpleBench | Grok 4 | xAI | 2025-07-09 | 0.61 |
| SimpleBench | Kimi K2 | Moonshot | 2025-07-12 | 0.26 |
| SimpleBench | GPT-5 | OpenAI | 2025-08-07 | 0.57 |
| SimpleBench | gpt-oss-120b | OpenAI | 2025-08-05 | 0.22 |
| SimpleBench | Claude Opus 4.1 | Anthropic | 2025-08-05 | 0.60 |
| FrontierMath | Grok 4 | xAI | 2025-07-09 | 0.12 |
| FrontierMath | GPT-4o | OpenAI | 2024-11-20 | 0.00 |
| FrontierMath | Mistral Large 2 | Mistral AI | 2024-11-18 | 0.00 |
| FrontierMath | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 0.03 |
| FrontierMath | o3-mini | OpenAI | 2025-01-31 | 0.08 |
| FrontierMath | Grok-2 | xAI | 2024-12-12 | 0.01 |
| FrontierMath | o3-mini | OpenAI | 2025-01-31 | 0.11 |
| FrontierMath | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 0.04 |
| FrontierMath | o1-mini | OpenAI | 2024-09-12 | 0.01 |
| FrontierMath | o1-mini | OpenAI | 2024-09-12 | 0.02 |
| FrontierMath | Claude 3.5 Sonnet | Anthropic | 2024-10-22 | 0.02 |
| FrontierMath | Claude 3.5 Haiku | Anthropic | 2024-10-22 | 0.00 |
| FrontierMath | Gemini 1.5 Flash | Google DeepMind | 2024-09-24 | 0.00 |
| FrontierMath | Claude 3.5 Sonnet | Anthropic | 2024-06-20 | 0.01 |
| FrontierMath | GPT-4o | OpenAI | 2024-08-06 | 0.00 |
| FrontierMath | o1 | OpenAI | 2024-12-17 | 0.09 |
| FrontierMath | DeepSeek-V3 | DeepSeek | 2024-12-26 | 0.02 |
| FrontierMath | Gemini 2.0 Flash | Google DeepMind,Google | 2025-02-05 | 0.02 |
| FrontierMath | Claude 3.7 Sonnet | Anthropic | 2025-02-24 | 0.03 |
| FrontierMath | Qwen2.5-Max | Alibaba | 2025-01-25 | 0.01 |
| FrontierMath | Llama 4 Scout | Meta AI | 2025-04-05 | 0.00 |
| FrontierMath | Llama 4 Maverick | Meta AI | 2025-04-05 | 0.01 |
| FrontierMath | Grok 3 | xAI | 2025-04-09 | 0.04 |
| FrontierMath | Grok-3 mini | xAI | 2025-04-09 | 0.06 |
| FrontierMath | Grok-3 mini | xAI | 2025-04-09 | 0.03 |
| FrontierMath | GPT-4.1 | OpenAI | 2025-04-14 | 0.06 |
| FrontierMath | GPT-4.1 mini | OpenAI | 2025-04-14 | 0.04 |
| FrontierMath | GPT-4.1 nano | OpenAI | 2025-04-14 | 0.01 |
| FrontierMath | o4-mini | OpenAI | 2025-04-16 | 0.17 |
| FrontierMath | o3 | OpenAI | 2025-04-16 | 0.10 |
| FrontierMath | o3 | OpenAI | 2025-04-16 | 0.10 |
| FrontierMath | o4-mini | OpenAI | 2025-04-16 | 0.10 |
| FrontierMath | o4-mini | OpenAI | 2025-04-16 | 0.19 |
| FrontierMath | Mistral Medium 3 | Mistral AI | 2025-05-07 | 0.00 |
| FrontierMath | Qwen Plus | Alibaba | 2025-04-28 | 0.02 |
| FrontierMath | Gemini 2.5 Pro | Google DeepMind | 2025-06-05 | 0.10 |
| FrontierMath | Gemini 2.5 Pro | Google DeepMind | 2025-06-17 | 0.11 |
| FrontierMath | Claude Sonnet 4 | Anthropic | 2025-05-22 | 0.04 |
| FrontierMath | Claude Opus 4 | Anthropic | 2025-05-22 | 0.04 |
| FrontierMath | Claude Opus 4.1 | Anthropic | 2025-08-05 | 0.06 |
| FrontierMath | Claude Opus 4.1 | Anthropic | 2025-08-05 | 0.07 |
| FrontierMath | Claude Opus 4 | Anthropic | 2025-05-22 | 0.04 |
| FrontierMath | Kimi K2 | Moonshot | 2025-07-12 | 0.02 |
| FrontierMath | GPT-5 nano | OpenAI | 2025-08-07 | 0.08 |
| FrontierMath | GPT-5 nano | OpenAI | 2025-08-07 | 0.07 |
| FrontierMath | GPT-5 | OpenAI | 2025-08-07 | 0.25 |
| FrontierMath | GPT-5 mini | OpenAI | 2025-08-07 | 0.19 |
| FrontierMath | GPT-5 mini | OpenAI | 2025-08-07 | 0.19 |
| FrontierMath | Claude Sonnet 4.5 | Anthropic | 2025-09-29 | 0.05 |
While these benchmarks do not capture all of the nuanced abilities needed for economically valuable tasks, the clear upward trends reflect genuine improvements in AI’s utility. This growth in capabilities shows no sign of slowing down.
Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.
Learn more about this graph
We show trends in a selection of benchmarks from Epoch’s Benchmarking Hub. These benchmarks are chosen to cover a diverse range of skills:
- FrontierMath tests models’ ability to solve research-level mathematical problems.
- GPQA Diamond poses graduate-level science questions across biology, physics, and chemistry, where answers are designed to be “Google-proof”.
- Aider Polyglot evaluates models’ performance on a set of challenging programming problems.
- SimpleBench is designed to test common sense reasoning by posing problems that are difficult for present-day models but easy for humans.
- The Visual Physics Comprehension Test (VPCT) is a benchmark designed to evaluate models’ understanding of basic physical scenarios.
Data
Data comes from a combination of internal evaluations and external reports. Aider Polyglot, SimpleBench, and VPCT are each collected from external benchmark leaderboards. GPQA Diamond and FrontierMath are run internally by Epoch. See the FAQs of our Benchmarking Hub for more information on how these evaluations are run.
Assumptions and limitations
The selected benchmarks cover only a subset of relevant skills which models have improved at. Notably, we do not yet track benchmarks for domains like robotics or biology, both of which appear to have improved notably over the past year.
Conversely, strong benchmark scores do not guarantee that models will generalize perfectly to real-world scenarios. For example, scoring 100% on GPQA Diamond does not immediately imply that a model can replace a human scientist; models can be overfit to benchmarks, and benchmarks typically do not capture all aspects of real-world work.


