OpenAI’s GPT-4 topped the Epoch Capabilities Index (ECI), our aggregate measure of language model capability, for roughly a year after its release in March 2023. No model since has led for as long. The second-longest lead, by OpenAI’s o1, lasted a little over three months, less than a third of GPT-4’s.
| Model | Model version ID | Display name | Organization | Country (of organization) | Version release date | Model accessibility | Publication date | Training compute (FLOP) | Confidence | Description | Unique display name | Model aggregation | Model group | Slug | ECI | ECI CI low | ECI CI high |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GLM-5.2 | glm-5.2_max | GLM-5.2 | Z.ai (Zhipu AI) | China | 2026-06-16 | Open weights (unrestricted) | 2026-06-16 | Unverified | GLM-5.2 (max) | GLM-5.2 | glm-5-2 | 152.72 | 149.02 | 159.70 | |||
| Qwen3.7-Max | qwen3.7-max | Alibaba | China | 2026-05-19 | API access | 2026-05-19 | Unverified | Qwen3.7-Max | qwen3-7-max | 153.19 | 150.76 | 160.42 | |||||
| DeepSeek-V4-Pro | deepseek-v4-pro_max | DeepSeek v4 Pro (max) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 9.7e+24 | Likely | DeepSeek v4 (max) | DeepSeek-V4-Pro | deepseek-v4-pro | 149.73 | 146.94 | 156.88 | ||
| Grok 4.3 Beta | grok-4.3_high | xAI | United States of America | 2026-04-17 | Hosted access (no API) | 2026-04-17 | Likely | Grok 4.3 Beta | grok-4-3-beta | 148.94 | 146.24 | 155.28 | |||||
| GPT-5.4 Pro | gpt-5.4-pro-2026-03-05_xhigh | GPT-5.4 Pro (xhigh) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 Pro (xhigh) | GPT-5.4 Pro | gpt-5-4-pro | 157.91 | 155.37 | 165.49 | |||
| GPT-5.2 Pro | gpt-5.2-pro-2025-12-11_xhigh | GPT-5.2 Pro | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Likely | GPT-5.2 Pro (xhigh) | GPT-5.2 Pro | gpt-5-2-pro | 154.72 | 152.20 | 161.41 | |||
| Kimi K2.7 Code | kimi-k2.7-code | Kimi K2.7 Code | Moonshot | China | 2026-06-12 | Open weights (unrestricted) | 2026-06-12 | Unverified | Kimi K2.7 Code | kimi-k2-7-code | 150.07 | 147.16 | 156.33 | ||||
| GPT-5.5 Pro | gpt-5.5-pro_xhigh | GPT-5.5 Pro (xhigh) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 Pro | gpt-5-5-pro | 161.29 | 158.43 | 170.30 | ||||
| GPT-5 Pro | gpt-5-pro-2025-10-06_high | GPT-5 Pro | OpenAI | United States of America | 2025-10-07 | API access | 2025-10-07 | Unknown | GPT-5 Pro (high) | GPT-5 Pro | gpt-5-pro | 150.48 | 147.41 | 157.34 | |||
| GPT-5 mini | gpt-5-mini-2025-08-07_high | GPT-5 mini (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 mini (high) | GPT-5 mini | gpt-5-mini | 145.83 | 143.09 | 151.48 | ||
| GPT-5.4 Nano | gpt-5.4-nano-2026-03-17_high | GPT-5.4 nano (high) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 nano (high) | GPT-5.4 Nano | gpt-5-4-nano | 146.54 | 143.59 | 152.29 | |||
| GPT-5 nano | gpt-5-nano-2025-08-07_high | GPT-5 nano (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 nano (high) | GPT-5 nano | gpt-5-nano | 140.85 | 136.04 | 146.03 | ||
| GPT-5.4 Mini | gpt-5.4-mini-2026-03-17_xhigh | GPT-5.4 mini (xhigh) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 mini (xhigh) | GPT-5.4 Mini | gpt-5-4-mini | 148.96 | 145.80 | 155.00 | |||
| AI Co-Mathematician | gdm-ai-co-mathematician | AI co-mathematician | Google DeepMind | United States of America | 2026-05-08 | Unreleased | 2026-05-08 | Unverified | AI Co-Mathematician | ai-co-mathematician | |||||||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-05 | Unknown | Google’s general-purpose frontier reasoning model. | Gemini 2.5 Pro | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.17 | 143.36 | 151.99 | ||
| Gemini 3 Flash | gemini-3-flash-preview | Google DeepMind | United States of America | 2025-12-17 | API access | 2025-12-17 | Unknown | Gemini 3 Flash | gemini-3-flash | 150.71 | 146.11 | 157.89 | |||||
| Gemini 3.1 Pro | gemini-3.1-pro-preview | Gemini 3.1 Pro Preview | Google DeepMind | United States of America | 2026-02-19 | API access | 2026-02-19 | Likely | Gemini 3.1 Pro | gemini-3-1-pro | 155.07 | 152.53 | 163.00 | ||||
| Claude Opus 4.1 | claude-opus-4-1-20250805_32K | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model, benchmarked with up to 32,000 reasoning tokens. | Claude Opus 4.1 (32k thinking) | Claude Opus 4.1 | claude-opus-4-1 | 144.34 | 141.71 | 150.46 | |||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_32K | Claude Sonnet 4.5 (32k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Anthropic's latest mid-size model, part of the Claude 4.5 series, evaluated with up to 32,000 reasoning tokens. | Claude Sonnet 4.5 (32k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 146.85 | 144.23 | 152.58 | ||
| GPT-5.4 | gpt-5.4-2026-03-05_xhigh | GPT-5.4 (xhigh) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (xhigh) | GPT-5.4 | gpt-5-4 | 156.22 | 153.71 | 163.79 | |||
| GPT-5.5 | gpt-5.5_xhigh | GPT-5.5 (xhigh) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.63 | 156.38 | 165.91 | ||||
| GPT-5.2 | gpt-5.2-2025-12-11_xhigh | GPT-5.2 (xhigh) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (xhigh) | GPT-5.2 | gpt-5-2 | 153.75 | 151.23 | 160.30 | |||
| o4-mini | o4-mini-2025-04-16_high | o4-mini (high) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | Unknown | The latest smaller size o-series reasoning model from OpenAI, evaluated with high reasoning effort. | o4-mini | o4-mini | 146.53 | 143.44 | 153.20 | |||
| o3-mini | o3-mini-2025-01-31_high | o3-mini (high) | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | Unknown | A smaller version of OpenAI's o3 reasoning model, evaluated with high reasoning effort. | o3-mini | o3-mini | 141.40 | 138.74 | 146.89 | |||
| Claude Opus 4.5 | claude-opus-4-5-20251101_32K | Claude Opus 4.5 (32k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (32k thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.62 | 146.60 | 155.89 | |||
| GPT-5 | gpt-5-2025-08-07_high | GPT-5 (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 (high) | GPT-5 | gpt-5 | 150.00 | 147.33 | 156.34 | |
| Claude Opus 4.6 | claude-opus-4-6_max | Claude Opus 4.6 (max) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (max) | Claude Opus 4.6 | claude-opus-4-6 | 155.55 | 152.30 | 163.09 | |||
| Claude Opus 4.7 | claude-opus-4-7_max | Claude Opus 4.7 (max) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (max) | Claude Opus 4.7 | claude-opus-4-7 | 156.31 | 153.46 | 163.87 | |||
| Claude Fable 5 | claude-fable-5_max | Claude Fable 5 (max) | Anthropic | United States of America | 2026-06-09 | API access | 2026-06-09 | Confident | Claude Fable 5 (max) | Claude Fable 5 | claude-fable-5 | 161.00 | 158.16 | 169.22 | |||
| Kimi K2.6 | kimi-k2.6 | Kimi K2.6 | Moonshot | China | 2026-04-20 | Open weights (unrestricted) | 2026-04-20 | Confident | Kimi K2.6 | kimi-k2-6 | 151.41 | 148.62 | 158.04 | ||||
| Gemini 3.5 Flash | gemini-3.5-flash_high | Gemini 3.5 Flash (high) | United States of America | 2026-05-19 | Unverified | Gemini 3.5 Flash | gemini-3-5-flash | 154.89 | 152.16 | 162.79 | |||||||
| Claude Opus 4.8 | claude-opus-4-8_max | Claude Opus 4.8 | Anthropic | United States of America | 2026-05-28 | API access | 2026-05-28 | Unverified | Claude Opus 4.8 (max) | Claude Opus 4.8 | claude-opus-4-8 | 158.27 | 154.62 | 165.90 | |||
| Claude Fable 5 | claude-fable-5_xhigh | Claude Fable 5 (xhigh) | Anthropic | United States of America | 2026-06-09 | API access | 2026-06-09 | Confident | Claude Fable 5 (xhigh) | Claude Fable 5 | claude-fable-5 | 161.00 | 158.16 | 169.22 | |||
| Qwen 3.6 Max (Preview) | qwen3.6-max-preview | Qwen 3.6 Max (Preview) | Alibaba | China | 2026-04-20 | Unverified | Qwen 3.6 Max (Preview) | qwen-3-6-max-preview | 149.78 | 147.09 | 157.28 | ||||||
| GLM-5.1 | glm-5.1 | GLM-5.1 | Z.ai (Zhipu AI) | China | 2026-04-07 | 2026-04-07 | Confident | GLM-5.1 | glm-5-1 | 150.67 | 147.27 | 156.75 | |||||
| Qwen 3.5 Plus (hosted 397B-A17B) | qwen3.5-plus | Alibaba | China | 2026-02-16 | 2026-02-16 | Likely | Qwen 3.5 Plus (hosted 397B-A17B) | qwen-3-5-plus-hosted-397b-a17b | 146.98 | 143.76 | 152.38 | ||||||
| Qwen 3.6 Plus | qwen3.6-plus | Qwen 3.6 Plus (2026-04-02) | Alibaba | China | 2026-03-31 | API access | 2026-04-01 | Confident | Qwen 3.6 Plus | qwen-3-6-plus | 149.40 | 145.69 | 156.20 | ||||
| Qwen 3.6 Flash | qwen3.6-flash | Qwen 3.6 Flash (2026-04-16) | Alibaba | China | 2026-04-27 | 2026-04-26 | Unverified | Qwen 3.6 Flash | qwen-3-6-flash | 144.75 | 141.26 | 150.11 | |||||
| Qwen 3.5 Flash (hosted 35B-A3B) | qwen3.5-flash | Alibaba | China | 2026-02-25 | 2026-02-25 | Likely | Qwen 3.5 Flash (hosted 35B-A3B) | qwen-3-5-flash-hosted-35b-a3b | 144.12 | 140.87 | 149.39 | ||||||
| GPT-5.5 | gpt-5.5_low | GPT-5.5 (low) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.63 | 156.38 | 165.91 | ||||
| GPT-5.5 Pro | gpt-5.5-pro-pre-release_xhigh | GPT-5.5 Pro (xhigh) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 Pro | gpt-5-5-pro | 161.29 | 158.43 | 170.30 | ||||
| GPT-5.5 | gpt-5.5-pre-release_xhigh | GPT-5.5 (xhigh) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.63 | 156.38 | 165.91 | ||||
| GPT-5.5 Pro | gpt-5.5-pro-pre-release_high | GPT-5.5 Pro (high) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 Pro | gpt-5-5-pro | 161.29 | 158.43 | 170.30 | ||||
| Claude Opus 4.7 | claude-opus-4-7_xhigh | Claude Opus 4.7 (xhigh) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (xhigh) | Claude Opus 4.7 | claude-opus-4-7 | 156.31 | 153.46 | 163.87 | |||
| GPT-5.4 Mini | gpt-5.4-mini-2026-03-17_high | GPT-5.4 mini (high) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 mini (high) | GPT-5.4 Mini | gpt-5-4-mini | 148.96 | 145.80 | 155.00 | |||
| Muse Spark | muse-spark | Muse Spark | Meta AI | United States of America | 2026-04-08 | API access | 2026-04-08 | Unverified | Muse Spark | Muse Spark | muse-spark | 154.57 | 151.11 | 162.37 | |||
| Gemma 4 31B IT | gemma-4-31b-it | Gemma 4 31B IT | Google DeepMind | United States of America | 2026-04-02 | Open weights (restricted use) | 2026-04-02 | Likely | Gemma 4 31B IT | Gemma 4 31B IT | gemma-4-31b-it | ||||||
| GPT-5.4 | gpt-5.4-2026-03-05_high | GPT-5.4 (high) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (high) | GPT-5.4 | gpt-5-4 | 156.22 | 153.71 | 163.79 | |||
| GPT-5.4 Pro | gpt-5.4-pro-2026-03-05-web-app | GPT-5.4 Pro (web) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 Pro (Web App) | GPT-5.4 Pro | gpt-5-4-pro | 157.91 | 155.37 | 165.49 | |||
| GPT-5.3 Codex | gpt-5.3-codex_high | GPT-5.3 Codex (high) | OpenAI | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | GPT-5.3 Codex | gpt-5-3-codex | 155.85 | 153.78 | 164.12 | ||||
| Gemini 3.1 Pro | gemini-3.1-pro-preview-customtools | Gemini 3.1 Pro Preview | Google DeepMind | United States of America | 2026-02-19 | API access | 2026-02-19 | Likely | Gemini 3.1 Pro | gemini-3-1-pro | 155.07 | 152.53 | 163.00 | ||||
| Claude Sonnet 4.6 | claude-sonnet-4-6_16K | Claude Sonnet 4.6 (16k thinking) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 153.41 | 150.31 | 160.16 | ||||
| Claude Sonnet 4.6 | claude-sonnet-4-6 | Claude Sonnet 4.6 | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 153.41 | 150.31 | 160.16 | ||||
| Claude Sonnet 4.6 | claude-sonnet-4-6_32K | Claude Sonnet 4.6 (32k thinking) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 153.41 | 150.31 | 160.16 | ||||
| Claude Opus 4.6 | claude-opus-4-6_120K | Claude Opus 4.6 (120k thinking) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (120k thinking) | Claude Opus 4.6 | claude-opus-4-6 | 155.55 | 152.30 | 163.09 | |||
| GLM-5 | glm-5 | GLM-5 | Z.ai (Zhipu AI) | China | 2026-02-11 | Open weights (unrestricted) | 2026-02-11 | Likely | GLM-5 | glm-5 | 146.50 | 143.80 | 152.25 | ||||
| Claude Opus 4.6 | claude-opus-4-6 | Claude Opus 4.6 (no thinking) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (no thinking) | Claude Opus 4.6 | claude-opus-4-6 | 155.55 | 152.30 | 163.09 | |||
| GPT-5.1 | gpt-5.1-2025-11-13_high | GPT-5.1 (high) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (high) | GPT-5.1 | gpt-5-1 | 149.84 | 147.18 | 156.15 | |||
| Kimi K2.5 | kimi-k2.5 | Kimi K2.5 | Moonshot | China | 2026-01-27 | Open weights (unrestricted) | 2026-02-02 | 5.8e+24 | Likely | Kimi K2.5 | kimi-k2-5 | 148.38 | 145.46 | 154.65 | |||
| Gemini 3 Pro | gemini-3-pro-preview | Gemini 3 Pro Preview | Google DeepMind | United States of America | 2025-11-18 | API access | 2025-11-18 | Unknown | Gemini 3 Pro Preview | Gemini 3 Pro | gemini-3-pro | 153.43 | 150.88 | 160.43 | |||
| GPT-5.2 | gpt-5.2-2025-12-11_high | GPT-5.2 (high) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (high) | GPT-5.2 | gpt-5-2 | 153.75 | 151.23 | 160.30 | |||
| o3 | o3-2025-04-16_medium | o3 (medium) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | Unknown | OpenAI's second generation o-series reasoning model, evaluated with medium reasoning effort. | o3 | o3 | 147.10 | 144.35 | 152.96 | |||
| Claude Opus 4.1 | claude-opus-4-1-20250805 | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model. | Claude Opus 4.1 | Claude Opus 4.1 | claude-opus-4-1 | 144.34 | 141.71 | 150.46 | |||
| GPT-4o | gpt-4o-2024-11-20 | GPT-4o (Nov 2024) | OpenAI | United States of America | 2024-11-20 | API access | 2024-05-13 | Speculative | The November 2024 version of GPT-4o, OpenAI's then-flagship multimodal language model. | GPT-4o (Nov 2024) | GPT-4o (Nov 2024) | gpt-4o-nov-2024 | 129.25 | 127.03 | 133.27 | ||
| GPT-4.1 | gpt-4.1-2025-04-14 | GPT-4.1 | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | Unknown | A coding-optimized GPT-4 series model from OpenAI. | GPT-4.1 | gpt-4-1 | 137.52 | 135.33 | 141.99 | |||
| Claude Opus 4.6 | claude-opus-4-6_64K | Claude Opus 4.6 (64k thinking) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (64k thinking) | Claude Opus 4.6 | claude-opus-4-6 | 155.55 | 152.30 | 163.09 | |||
| Claude Opus 4.6 | claude-opus-4-6_32K | Claude Opus 4.6 (32k thinking) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (32k thinking) | Claude Opus 4.6 | claude-opus-4-6 | 155.55 | 152.30 | 163.09 | |||
| Claude Opus 4 | claude-opus-4-20250514 | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1). | Claude Opus 4 | Claude Opus 4 | claude-opus-4 | 143.10 | 140.42 | 148.78 | ||
| Claude Opus 4.5 | claude-opus-4-5-20251101 | Claude Opus 4.5 (no thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (no thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.62 | 146.60 | 155.89 | |||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929 | Claude Sonnet 4.5 (no thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Anthropic's latest mid-size model, part of the Claude 4.5 series. | Claude Sonnet 4.5 (no thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 146.85 | 144.23 | 152.58 | ||
| GPT-5 | gpt-5-2025-08-07_medium | GPT-5 (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 (medium) | GPT-5 | gpt-5 | 150.00 | 147.33 | 156.34 | |
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219 | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic. | Claude 3.7 Sonnet | Claude 3.7 Sonnet | claude-3-7-sonnet | 141.83 | 139.34 | 147.11 | ||
| Kimi K2.5 | fireworks/kimi-k2p5 | Kimi K2.5 (Fireworks) | Moonshot | China | 2026-01-27 | Open weights (unrestricted) | 2026-02-02 | 5.8e+24 | Likely | Kimi K2.5 | kimi-k2-5 | 148.38 | 145.46 | 154.65 | |||
| GPT-5 mini | gpt-5-mini-2025-08-07_medium | GPT-5 mini (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 mini (medium) | GPT-5 mini | gpt-5-mini | 145.83 | 143.09 | 151.48 | ||
| Grok 4 | grok-4-0709 | xAI | United States of America | 2025-07-09 | API access | 2025-07-09 | 5.0e+26 | Speculative | Grok 4 | Grok 4 | grok-4 | 147.10 | 144.69 | 153.68 | |||
| GLM-4.7 | zai-org/GLM-4.7 | GLM-4.7 (Together) | Z.ai (Zhipu AI) | China | 2025-12-22 | Open weights (unrestricted) | 2025-12-22 | 4.4e+24 | Likely | GLM-4.7 | glm-4-7 | 143.84 | 139.56 | 149.42 | |||
| GLM-4.7 | glm-4.7 | GLM-4.7 | Z.ai (Zhipu AI) | China | 2025-12-22 | Open weights (unrestricted) | 2025-12-22 | 4.4e+24 | Likely | GLM-4.7 | GLM-4.7 | glm-4-7 | 143.84 | 139.56 | 149.42 | ||
| GPT-5.2 Pro | gpt-5.2-pro-2025-12-11-webapp | GPT-5.2 Pro (web) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Likely | GPT-5.2 Pro (Web App) | GPT-5.2 Pro | gpt-5-2-pro | 154.72 | 152.20 | 161.41 | |||
| DeepSeek-V3.2 | fireworks/deepseek-v3p2 | DeepSeek-V3.2 (Thinking; Fireworks) | DeepSeek | China | 2025-12-01 | Open weights (unrestricted) | 2025-12-01 | 4.2e+24 | Likely | DeepSeek-V3.2 | deepseek-v3-2 | 146.64 | 143.60 | 152.74 | |||
| Gemini 2.5 Flash | gemini-2.5-flash | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-04-17 | Unknown | A small model from Google DeepMind's Gemini 2.5 series. | Gemini 2.5 Flash (Jun 2025) | Gemini 2.5 Flash (Jun 2025) | Gemini 2.5 Flash (Jun 2025) | gemini-2-5-flash-jun-2025 | 140.31 | 137.58 | 145.90 | ||
| DeepSeek-V3.2 | deepseek-reasoner | DeepSeek-V3.2 (Thinking) | DeepSeek | China | 2025-12-01 | Open weights (unrestricted) | 2025-12-01 | 4.2e+24 | Likely | DeepSeek-V3.2 | deepseek-v3-2 | 146.64 | 143.60 | 152.74 | |||
| gpt-oss-120b | openai/gpt-oss-120b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | gpt-oss-120b (high) | gpt-oss-120b | gpt-oss-120b | 140.55 | 135.47 | 146.61 | |||
| Qwen3-235B-A22B (Jul 2025) | Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | A 22 billion active (235 billion total)-parameter mixture-of-experts reasoning model in the Qwen 3 series. | Qwen3-235B-A22B-Thinking (Jul 2025) | Qwen3-235B-A22B-Thinking (Jul 2025) | qwen3-235b-a22b-thinking-jul-2025 | 144.95 | 141.90 | 151.93 | ||
| GPT-5.2 | gpt-5.2-2025-12-11_medium | GPT-5.2 (medium) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (medium) | GPT-5.2 | gpt-5-2 | 153.75 | 151.23 | 160.30 | |||
| GPT-5.2 | gpt-5.2-2025-12-11_low | GPT-5.2 (low) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (low) | GPT-5.2 | gpt-5-2 | 153.75 | 151.23 | 160.30 | |||
| Qwen3-235B-A22B (Jul 2025) | Qwen/Qwen3-235B-A22B-Thinking-2507 | Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | Qwen3-235B-A22B-Thinking (Jul 2025) | Qwen3-235B-A22B-Thinking (Jul 2025) | qwen3-235b-a22b-thinking-jul-2025 | 144.95 | 141.90 | 151.93 | ||
| Qwen3-Max | qwen3-max-2025-09-23 | Qwen3-Max-Instruct | Alibaba | China | 2025-09-24 | API access | 2025-09-05 | 1.5e+25 | Speculative | A 1-trillion total parameter scale model in the Qwen 3 series. | Qwen3 Max | Qwen3-Max | qwen3-max | 144.24 | 140.46 | 159.83 | |
| Kimi K2 Thinking | kimi-k2-thinking-turbo | Kimi K2 Thinking Turbo | Moonshot | China | 2025-11-06 | Open weights (restricted use) | 2025-11-06 | 4.2e+24 | Likely | Kimi K2 Thinking | kimi-k2-thinking | 146.08 | 142.98 | 151.63 | |||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_59K | Claude Sonnet 4.5 (59k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (59k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 146.85 | 144.23 | 152.58 | |||
| Claude Opus 4.1 | claude-opus-4-1-20250805_27K | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model, benchmarked with up to 27,000 reasoning tokens. | Claude Opus 4.1 (27k thinking) | Claude Opus 4.1 | claude-opus-4-1 | 144.34 | 141.71 | 150.46 | |||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_32K | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (32k thinking) | Claude Haiku 4.5 | claude-haiku-4-5 | 142.91 | 139.17 | 148.79 | ||||
| o3 | o3-2025-04-16_high | o3 (high) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | Unknown | OpenAI's second generation o-series reasoning model, evaluated with high reasoning effort. | o3 | o3 | 147.10 | 144.35 | 152.96 | |||
| Grok-3 mini | grok-3-mini-beta_high | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | Unknown | A beta release of a smaller version of XAI's third generation Grok model, evaluated with high reasoning effort. | Grok-3 mini | grok-3-mini | 141.05 | 138.50 | 145.60 | ||||
| Kimi K2 Thinking | moonshotai/Kimi-K2-Thinking | Kimi K2 Thinking (Together) | Moonshot | China | 2025-11-06 | Open weights (restricted use) | 2025-11-06 | 4.2e+24 | Likely | Kimi K2 Thinking | kimi-k2-thinking | 146.08 | 142.98 | 151.63 | |||
| GLM-4.6 | zai-org/GLM-4.6 | GLM-4.6 (Together) | Z.ai (Zhipu AI),Tsinghua University | China | 2025-09-30 | Open weights (unrestricted) | 2025-09-30 | 4.4e+24 | Likely | GLM-4.6 | glm-4-6 | 140.96 | 132.50 | 147.47 | |||
| DeepSeek-R1 (May 2025) | DeepSeek-R1-0528 | DeepSeek-R1 (May 2025) | DeepSeek | China | 2025-05-28 | Open weights (unrestricted) | 2025-05-28 | 4.0e+24 | Confident | An updated May 2025 version of DeepSeek's reasoning model, R1. | DeepSeek-R1 (May 2025) | DeepSeek-R1 (May 2025) | deepseek-r1-may-2025 | 142.00 | 139.40 | 147.19 | |
| Claude 3.5 Haiku | claude-3-5-haiku-20241022 | Claude 3.5 Haiku (Oct 2024) | Anthropic | United States of America | 2024-10-22 | API access | 2024-10-22 | Unknown | The smallest model in Anthropic’s Claude 3.5 Series. | Claude 3.5 Haiku (Oct 2024) | Claude 3.5 Haiku | claude-3-5-haiku | 127.45 | 123.20 | 132.02 | ||
| GPT-5.1 | gpt-5.1-2025-11-13_none | GPT-5.1 (no thinking) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (no thinking) | GPT-5.1 | gpt-5-1 | 149.84 | 147.18 | 156.15 | |||
| GPT-5.1 | gpt-5.1-2025-11-13_low | GPT-5.1 (low) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (low) | GPT-5.1 | gpt-5-1 | 149.84 | 147.18 | 156.15 | |||
| Claude Opus 4.5 | claude-opus-4-5-20251101_16K | Claude Opus 4.5 (16k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (16k thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.62 | 146.60 | 155.89 | |||
| GPT-5.1 | gpt-5.1-2025-11-13_medium | GPT-5.1 (medium) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (medium) | GPT-5.1 | gpt-5-1 | 149.84 | 147.18 | 156.15 | |||
| o3 | o3-2025-04-16_low | o3 (low) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | Unknown | OpenAI's latest full-size o-series reasoning model release, evaluated on low reasoning effort. | o3 | o3 | 147.10 | 144.35 | 152.96 | |||
| o4-mini | o4-mini-2025-04-16_low | o4-mini (low) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | Unknown | The latest smaller size o-series reasoning model from OpenAI, evaluated with low reasoning effort. | o4-mini | o4-mini | 146.53 | 143.44 | 153.20 | |||
| o4-mini | o4-mini-2025-04-16_medium | o4-mini (medium) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | Unknown | The latest smaller size o-series reasoning model from OpenAI, evaluated with medium reasoning effort. | o4-mini | o4-mini | 146.53 | 143.44 | 153.20 | |||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_16K | Claude Sonnet 4.5 (16k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Anthropic's latest mid-size model, part of the Claude 4.5 series, evaluated with up to 16,000 reasoning tokens. | Claude Sonnet 4.5 (16k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 146.85 | 144.23 | 152.58 | ||
| GPT-4 (Jun 2023) | gpt-4-0613 | GPT-4 (Jun 2023) | OpenAI | United States of America | 2023-06-13 | API access | 2023-06-13 | 2.1e+25 | Likely | The June 2023 version of GPT-4, OpenAI's 2023 multimodal flagship model. | GPT-4 (Jun 2023) | gpt-4-jun-2023 | 121.73 | 118.41 | 125.17 | ||
| GPT-4 (Mar 2023) | gpt-4-0314 | GPT-4 (Mar 2023) | OpenAI | United States of America | 2023-03-14 | API access | 2023-03-15 | 2.1e+25 | Likely | The March 2023 version of GPT-4, OpenAI's 2023 multimodal flagship model. | GPT-4 (Mar 2023) | gpt-4-mar-2023 | 125.82 | 121.38 | 133.35 | ||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001 | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 | claude-haiku-4-5 | 142.91 | 139.17 | 148.79 | |||||
| Grok 4 Heavy | grok-4-heavy-web-app | Grok 4 Heavy (web app) | xAI | United States of America | 2025-07-10 | Unreleased | 2025-07-10 | Unknown | Grok 4 Heavy | grok-4-heavy | |||||||
| GLM-4.5 | glm-4.5 | Z.ai (Zhipu AI),Tsinghua University | China | 2025-08-03 | Open weights (unrestricted) | 2025-08-05 | 4.4e+24 | Confident | Zhipu AI & Tsinghua University’s large open source model. | GLM-4.5 | glm-4-5 | ||||||
| GPT-5 nano | gpt-5-nano-2025-08-07_medium | GPT-5 nano (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 nano (medium) | GPT-5 nano | gpt-5-nano | 140.85 | 136.04 | 146.03 | ||
| Claude Opus 4 | claude-opus-4-20250514_27K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 27,000 reasoning tokens. | Claude Opus 4 (27k thinking) | Claude Opus 4 | claude-opus-4 | 143.10 | 140.42 | 148.78 | ||
| Claude Opus 4.1 | claude-opus-4-1-20250805_16K | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model, benchmarked with up to 16,000 reasoning tokens. | Claude Opus 4.1 (16k thinking) | Claude Opus 4.1 | claude-opus-4-1 | 144.34 | 141.71 | 150.46 | |||
| GPT-4o mini | gpt-4o-mini-2024-07-18 | OpenAI | United States of America | 2024-07-18 | API access | 2024-07-18 | Speculative | A smaller version of OpenAI's GPT-4o. | GPT-4o mini | gpt-4o-mini | 126.87 | 122.15 | 129.93 | ||||
| Claude Sonnet 4 | claude-sonnet-4-20250514 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized version of Claude 4 series of reasoning models. | Claude Sonnet 4 | Claude Sonnet 4 | claude-sonnet-4 | 142.32 | 139.26 | 147.36 | |||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro-preview-06-05 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America | 2025-06-05 | API access | 2025-06-05 | Unknown | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model. | Gemini 2.5 Pro Preview (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.17 | 143.36 | 151.99 | |
| Claude 3.5 Sonnet (October 2024) | claude-3-5-sonnet-20241022 | Claude 3.5 Sonnet (Oct 2024) | Anthropic | United States of America | 2024-10-22 | API access | 2024-10-22 | Unverified | An updated version of the top model in the Claude 3.5 series, a previous flagship model series from Anthropic. | Claude 3.5 Sonnet (Oct 2024) | Claude 3.5 Sonnet (October 2024) | claude-3-5-sonnet-october-2024 | 134.34 | 130.85 | 141.01 | ||
| Claude 3.5 Sonnet | claude-3-5-sonnet-20240620 | Claude 3.5 Sonnet (Jun 2024) | Anthropic | United States of America | 2024-06-20 | API access | 2024-06-20 | 2.7e+25 | Speculative | The first version of the top model in the Claude 3.5 series, a previous flagship model series from Anthropic. | Claude 3.5 Sonnet (Jun 2024) | Claude 3.5 Sonnet | claude-3-5-sonnet | 130.00 | 127.27 | 137.45 | |
| Claude Sonnet 4 | claude-sonnet-4-20250514_59K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 59,000 reasoning tokens. | Claude Sonnet 4 (59k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.32 | 139.26 | 147.36 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_64K | Claude 3.7 Sonnet (64k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 64,000 reasoning tokens allowed. | Claude 3.7 Sonnet (64k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 141.83 | 139.34 | 147.11 | |
| Grok 3 | grok-3-beta | xAI | United States of America | 2025-04-09 | API access | 2025-02-17 | 3.5e+26 | Likely | A beta release of XAI's third generation Grok model. | Grok 3 (beta) | Grok 3 | grok-3 | 139.06 | 136.97 | 143.60 | ||
| Magistral Small 1.1 | magistral-small-2506 | Mistral AI | France | 2025-06-10 | Open weights (unrestricted) | 2025-06-10 | Confident | Magistral Small 1.1 | magistral-small-1-1 | 133.05 | 127.70 | 137.17 | |||||
| Qwen3-235B-A22B | qwen3-235b-a22b | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Likely | Qwen3 (Apr 2025) | Qwen3-235B-A22B | qwen3-235b-a22b | 139.61 | 134.81 | 144.40 | |||
| Gemini 2.5 Pro (May 2025) | gemini-2.5-pro-preview-05-06 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America | 2025-05-06 | API access | 2025-05-06 | Unknown | A May 2025 preview version of Gemini 2.5 Pro, a flagship reasoning model from Google DeepMind. | Gemini 2.5 Pro Preview (Jun 2025) | Gemini 2.5 Pro (May 2025) | Gemini 2.5 Pro (May 2025) | gemini-2-5-pro-may-2025 | 142.60 | 139.97 | 148.49 | |
| DeepSeek-R1 | DeepSeek-R1 | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-20 | 3.5e+24 | Confident | DeepSeek-R1 | DeepSeek-R1 | deepseek-r1 | 139.71 | 137.57 | 144.43 | |||
| Claude Sonnet 4 | claude-sonnet-4-20250514_32K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 32,000 reasoning tokens. | Claude Sonnet 4 (32k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.32 | 139.26 | 147.36 | |||
| Claude Opus 4 | claude-opus-4-20250514_16K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 16,000 reasoning tokens. | Claude Opus 4 (16k thinking) | Claude Opus 4 | claude-opus-4 | 143.10 | 140.42 | 148.78 | ||
| Claude Sonnet 4 | claude-sonnet-4-20250514_16K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 16,000 reasoning tokens. | Claude Sonnet 4 (16k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.32 | 139.26 | 147.36 | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20 | Google DeepMind | United States of America | 2025-05-20 | API access | 2025-04-17 | Unknown | The May 2025 preview version of Gemini 2.5 Flash, a small reasoning model from Google DeepMind's Gemini 2.5 series. | Gemini 2.5 Flash (May 2025) | Gemini 2.5 Flash (May 2025) | gemini-2-5-flash-may-2025 | 142.30 | 139.70 | 147.65 | |||
| Qwen Plus | qwen-plus-2025-04-28 | Alibaba | China | 2025-04-28 | API access | 2024-02-06 | Unknown | Qwen Plus (Apr 2025) | Qwen Plus (Apr 2025) | Qwen Plus (Apr 2025) | qwen-plus-apr-2025 | ||||||
| DeepSeek-V3 (Mar 2025) | DeepSeek-V3-0324 | DeepSeek-V3 (Mar 2025) | DeepSeek | China | 2025-03-24 | Open weights (restricted use) | 2025-03-24 | 3.3e+24 | Confident | An updated version of the DeepSeek V3 model. | DeepSeek-V3 (Mar 2025) | DeepSeek-V3 (Mar 2025) | deepseek-v3-mar-2025 | 137.07 | 134.38 | 142.52 | |
| Gemini 2.5 Pro (Mar 2025) | gemini-2.5-pro-preview-03-25 | Gemini 2.5 Pro Preview (Mar 2025) | Google DeepMind | United States of America | 2025-03-31 | API access | 2025-03-25 | Unknown | A March 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship model. | Gemini 2.5 Pro Preview (Mar 2025) | Gemini 2.5 Pro (Mar 2025) | Gemini 2.5 Pro (Mar 2025) | gemini-2-5-pro-mar-2025 | 144.92 | 141.98 | 151.59 | |
| Mistral Medium 3 | mistral-medium-2505 | Mistral AI | France | 2025-05-07 | API access | 2025-05-07 | Unknown | A 2025 language model from Mistral. | Mistral Medium 3 | mistral-medium-3 | 135.29 | 132.22 | 139.52 | ||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-04-17 | Gemini 2.5 Flash Preview (Apr 2025) | Google DeepMind | United States of America | 2025-04-17 | API access | 2025-04-17 | Unknown | An April 2025 preview of Gemini 2.5 Flash, a small model in Google DeepMind's Gemini 2.5 series. | Gemini 2.5 Flash (Apr 2025) | Gemini 2.5 Flash (Apr 2025) | gemini-2-5-flash-apr-2025 | 140.55 | 136.89 | 146.50 | ||
| GPT-4.1 nano | gpt-4.1-nano-2025-04-14 | GPT-4.1 nano | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | Unknown | The smallest version of OpenAI's GPT-4.1. | GPT-4.1 nano | gpt-4-1-nano | 130.85 | 126.13 | 135.91 | |||
| GPT-4.1 mini | gpt-4.1-mini-2025-04-14 | GPT-4.1 mini | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | Unknown | A smaller version of OpenAI's GPT-4.1. | GPT-4.1 mini | gpt-4-1-mini | 135.79 | 132.71 | 140.13 | |||
| QWQ-Plus | qwq-plus | Alibaba | China | 2025-04-08 | API access | 2025-04-08 | Unknown | QwQ-Plus | QWQ-Plus | qwq-plus | |||||||
| Grok-3 mini | grok-3-mini-beta_low | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | Unknown | A beta release of a smaller version of XAI's third generation Grok model, evaluated with low reasoning effort. | Grok-3 mini | grok-3-mini | 141.05 | 138.50 | 145.60 | ||||
| Llama 4 Maverick | Llama-4-Maverick-17B-128E-Instruct-FP8 | Llama 4 Maverick (FP8) | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | Likely | The 17 billion active (400 billion total)-parameter model in Meta's Llama 4 series, quantized to FP8. | Llama 4 Maverick | llama-4-maverick | 133.04 | 130.31 | 137.25 | ||
| Llama 4 Scout | Llama-4-Scout-17B-16E-Instruct | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 4.1e+24 | Likely | The 17 billion active (109 billion total)-parameter model in Meta's Llama 4 series. | Llama 4 Scout | llama-4-scout | 130.52 | 127.97 | 134.52 | |||
| Qwen-Turbo | qwen-turbo-2024-11-01 | Alibaba | China | 2024-11-01 | API access | 2024-02-06 | Unknown | Qwen Turbo | Qwen-Turbo | qwen-turbo | |||||||
| Qwen Plus | qwen-plus-2025-01-25 | Alibaba | China | 2025-01-25 | API access | 2024-02-06 | Unknown | Qwen Plus (Jan 2025) | Qwen Plus (Jan 2025) | Qwen Plus (Jan 2025) | qwen-plus-jan-2025 | ||||||
| Qwen2.5-Max | qwen-max-2025-01-25 | Alibaba | China | 2025-01-25 | API access | 2025-01-28 | Unknown | Qwen2.5-Max | Qwen2.5-Max | qwen2-5-max | 133.41 | 130.61 | 139.10 | ||||
| Hermes 2 Theta Llama-3 70B | Hermes-2-Theta-Llama-3-70B | Nous Research,Arcee AI | United States of America | 2024-06-20 | Open weights (restricted use) | 2024-06-20 | Confident | Nous Research’s fine-tuned Llama-3 70B optimized for instruction following and chat. | Hermes 2 Theta Llama-3 70B | hermes-2-theta-llama-3-70b | |||||||
| Gemini 2.5 Pro (Mar 2025) | gemini-2.5-pro-exp-03-25 | Gemini 2.5 Pro Exp (Mar 2025) | Google DeepMind | United States of America | 2025-03-25 | API access | 2025-03-25 | Unknown | A March 2025 preview version of Gemini 2.5 Pro. | Gemini 2.5 Pro Exp (Mar 2025) | Gemini 2.5 Pro (Mar 2025) | Gemini 2.5 Pro (Mar 2025) | gemini-2-5-pro-mar-2025 | 144.92 | 141.98 | 151.59 | |
| Mistral Small 3 | mistral-small-2501 | Mistral AI | France | 2025-01-25 | Open weights (unrestricted) | 2025-01-30 | 1.2e+24 | Confident | Mistral Small 3 | mistral-small-3 | |||||||
| Mistral Small 3.1 | mistral-small-2503 | Mistral AI | France | 2025-03-17 | Open weights (unrestricted) | 2025-03-17 | Confident | An updated version of Mistral small, a 24 billion-parameter model from Mistral. | Mistral Small 3.1 | mistral-small-3-1 | |||||||
| Gemma 3 27B | gemma-3-27b-it | Google DeepMind | United States of America | 2025-03-12 | Open weights (restricted use) | 2025-03-12 | 2.3e+24 | Confident | An instruction-tuned version of the 27 billion-parameter model Google DeepMind's Gemma 3 series. | Gemma 3 27B | gemma-3-27b | 131.05 | 127.90 | 135.66 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_32K | Claude 3.7 Sonnet (32k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 32,000 reasoning tokens allowed. | Claude 3.7 Sonnet (32k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 141.83 | 139.34 | 147.11 | |
| Gemini 1.5 Flash 8B | gemini-1.5-flash-8b-001 | Google DeepMind | United States of America | 2024-10-03 | API access | 2024-05-10 | Confident | Gemini 1.5 Flash 8B | gemini-1-5-flash-8b | ||||||||
| DeepSeek-R1-Distill-Qwen-14B | DeepSeek-R1-Distill-Qwen-14B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | Speculative | A model based on Qwen 14B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | DeepSeek-R1-Distill-Qwen-14B | deepseek-r1-distill-qwen-14b | |||||||
| DeepSeek-R1-Distill-Llama-70B | DeepSeek-R1-Distill-Llama-70B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | Speculative | A model based on LLaMA 70B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | DeepSeek-R1-Distill-Llama-70B | deepseek-r1-distill-llama-70b | |||||||
| Gemini 2.0 Flash | gemini-2.0-flash-001 | Gemini 2.0 Flash (Feb 2025) | Google DeepMind,Google | United States of America | 2025-02-05 | API access | 2024-12-11 | Unknown | The first version of Google DeepMind's Gemini 2.0 Flash model. | Gemini 2.0 Flash (Feb 2025) | Gemini 2.0 Flash (Feb 2025) | gemini-2-0-flash-feb-2025 | 135.80 | 133.35 | 140.78 | ||
| DeepSeek-V3 | DeepSeek-V3 | DeepSeek | China | 2024-12-26 | Open weights (restricted use) | 2024-12-24 | 3.3e+24 | Confident | DeepSeek’s 2024 mixture-of-experts model. | DeepSeek-V3 | DeepSeek-V3 | deepseek-v3 | 133.03 | 129.36 | 139.00 | ||
| Tulu 3 (Tülu 3) 70B | Llama-3.1-Tulu-3-70B-DPO | Allen Institute for AI,University of Washington | United States of America | 2024-11-21 | Open weights (restricted use) | 2024-11-21 | Confident | Tülu 3 70B | Tulu 3 (Tülu 3) 70B | tulu-3-tlu-3-70b | |||||||
| Gemma 2 27B | gemma-2-27b-it | Google DeepMind | United States of America | 2024-06-24 | Open weights (restricted use) | 2024-06-24 | 2.1e+24 | Confident | An instruction-optimized version of Google DeepMind's Gemma 2 27B. | Gemma 2 27B | gemma-2-27b | 122.32 | 119.43 | 125.26 | |||
| Gemma 2 9B | gemma-2-9b-it | Google DeepMind | United States of America | 2024-06-24 | Open weights (restricted use) | 2024-06-24 | 4.3e+23 | Confident | Gemma 2 9B | gemma-2-9b | 119.35 | 115.40 | 122.93 | ||||
| Claude 2.1 | claude-2.1 | Anthropic | United States of America | 2023-11-21 | API access | 2023-11-21 | Unknown | An updated version in Anthropic's Claude 2 series. | Claude 2.1 | claude-2-1 | 117.89 | 108.94 | 120.75 | ||||
| Claude 2 | claude-2.0 | Anthropic | United States of America | 2023-07-11 | API access | 2023-07-11 | 3.9e+24 | Speculative | Anthropic's second generation Claude model. | Claude 2 | claude-2 | 119.36 | 115.73 | 125.35 | |||
| Gemini 2.0 Flash Thinking | gemini-2.0-flash-thinking-exp-01-21 | Gemini 2.0 Flash Thinking Exp | Google DeepMind,Google | United States of America | 2025-01-21 | API access | 2024-12-19 | Unknown | A January 2025 experimental version of Gemini 2.0 Flash Thinking, a small reasoning model from Google DeepMind. | Gemini 2.0 Flash Thinking (Jan 2025) | Gemini 2.0 Flash Thinking (Jan 2025) | gemini-2-0-flash-thinking-jan-2025 | 136.25 | 131.42 | 141.68 | ||
| o1-preview | o1-preview-2024-09-12 | o1-preview | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | Unknown | The September 2024 preview version of OpenAI’s first reasoning model, o1. | o1-preview | o1-preview | 135.85 | 132.37 | 141.76 | |||
| Gemini 2.0 Pro | gemini-2.0-pro-exp-02-05 | Gemini 2.0 Pro Exp (Feb 2025) | Google DeepMind | United States of America | 2025-02-05 | Hosted access (no API) | 2024-12-11 | Unknown | A February 2025 experimental version of Google DeepMind's previous flagship model, Gemini 2.0 Pro. | Gemini 2.0 Pro | gemini-2-0-pro | 135.63 | 131.98 | 140.02 | |||
| o1 | o1-2024-12-17_high | o1 (high) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | Unknown | OpenAI's first reasoning model, evaluated with at the high reasoning effort level. | o1 | o1 | 142.74 | 140.13 | 147.88 | |||
| GPT-4o | gpt-4o-2024-08-06 | GPT-4o (Aug 2024) | OpenAI | United States of America | 2024-08-06 | API access | 2024-05-13 | Speculative | An updated version of GPT-4o, OpenAI's model that powered ChatGPT from mid-2024 to -2025. | GPT-4o (Aug 2024) | GPT-4o (Aug 2024) | gpt-4o-aug-2024 | 129.14 | 126.45 | 133.08 | ||
| Gemini 1.5 Flash | gemini-1.5-flash-002 | Google DeepMind | United States of America | 2024-09-24 | API access | 2024-05-10 | Unknown | The second version of Google DeepMind's Gemini 1.5 Flash. | Gemini 1.5 Flash (Sep 2024) | Gemini 1.5 Flash (Sep 2024) | gemini-1-5-flash-sep-2024 | 130.35 | 126.39 | 134.60 | |||
| o1-mini | o1-mini-2024-09-12_high | o1-mini (high) | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | Unknown | A smaller version of OpenAI’s first reasoning model, o1, evaluated with high reasoning effort. | o1-mini | o1-mini | 136.56 | 133.72 | 141.20 | |||
| o1-mini | o1-mini-2024-09-12_medium | o1-mini (medium) | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | Unknown | A smaller version of OpenAI’s first reasoning model, o1, evaluated with medium reasoning effort. | o1-mini | o1-mini | 136.56 | 133.72 | 141.20 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_16K | Claude 3.7 Sonnet (16k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 16,000 reasoning tokens allowed. | Claude 3.7 Sonnet (16k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 141.83 | 139.34 | 147.11 | |
| o3-mini | o3-mini-2025-01-31_medium | o3-mini (medium) | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | Unknown | A smaller version of OpenAI's o3 reasoning model, evaluated with medium reasoning effort. | o3-mini | o3-mini | 141.40 | 138.74 | 146.89 | |||
| Grok-2 | grok-2-1212 | xAI | United States of America | 2024-12-12 | API access | 2024-08-13 | 3.0e+25 | Confident | XAI's second generation Grok model. | Grok-2 (Dec 2024) | Grok-2 (Dec 2024) | grok-2-dec-2024 | 130.76 | 128.30 | 134.60 | ||
| Mistral Large 2 | mistral-large-2411 | Mistral AI | France | 2024-11-18 | Open weights (non-commercial) | 2024-07-24 | 2.1e+25 | Likely | Mistral Large 2 (Nov 2024) | Mistral Large 2 (Nov 2024) | mistral-large-2-nov-2024 | 128.74 | 126.10 | 132.83 | |||
| GPT-4.5 | gpt-4.5-preview-2025-02-27 | GPT-4.5 Preview (Feb 2025) | OpenAI | United States of America | 2025-02-27 | API access | 2025-02-27 | 3.8e+26 | Likely | The largest model in OpenAI’s GPT series. | GPT-4.5 | gpt-4-5 | 137.60 | 135.29 | 143.04 | ||
| GPT-4 Turbo (Apr 2024) | gpt-4-turbo-2024-04-09 | OpenAI | United States of America | 2024-04-09 | API access | 2024-04-09 | Unknown | The April 2024 version of GPT-4 Turbo, OpenAI's then-flagship language model. | GPT-4 Turbo (Apr 2024) | gpt-4-turbo-apr-2024 | 127.57 | 124.94 | 131.48 | ||||
| o1 | o1-2024-12-17_medium | o1 (medium) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | Unknown | OpenAI's first reasoning model, evaluated with at the medium reasoning effort level. | o1 | o1 | 142.74 | 140.13 | 147.88 | |||
| Llama 3-70B | Meta-Llama-3-70B-Instruct | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.9e+24 | Confident | An instruction-tuned version of the 70 billion-parameter model in Meta’s LLaMA 3 series. | Llama 3-70B | llama-3-70b | 122.47 | 120.41 | 126.46 | |||
| Gemini 1.5 Pro | gemini-1.5-pro-001 | Google DeepMind | United States of America | 2024-05-14 | API access | 2024-02-15 | Speculative | The first of two versions of a previous flagship model from Google DeepMind, Gemini 1.5 Pro. | Gemini 1.5 Pro (May 2024) | Gemini 1.5 Pro (May 2024) | gemini-1-5-pro-may-2024 | 127.12 | 124.27 | 131.65 | |||
| Llama 3.3 70B | Llama-3.3-70B-Instruct | Meta AI | United States of America | 2024-12-06 | Open weights (restricted use) | 2024-12-06 | 6.9e+24 | Confident | An instruction-tuned version of the 70 billion-parameter model in Meta's Llama 3.3 series. | Llama 3.3 70B | llama-3-3-70b | 127.46 | 125.24 | 131.89 | |||
| Llama 3.2 90B | Llama-3.2-90B-Vision-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | Confident | An instruction-tuned version of the 90 billion-parameter vision model in Meta's Llama 3.3 series. | Llama 3.2 90B | llama-3-2-90b | 125.61 | 122.63 | 129.34 | ||||
| Llama 2-70B | Llama-2-70b-chat-hf | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | Confident | A chat-optimized version of the 70-billion parameter model in Meta's Llama 2 series. | Llama 2-70B | llama-2-70b | 113.21 | 109.61 | 116.25 | |||
| Gemini 1.5 Pro | gemini-1.5-pro-002 | Google DeepMind | United States of America | 2024-09-24 | API access | 2024-02-15 | Speculative | The first of two versions of a previous flagship model from Google DeepMind, Gemini 1.5 Pro. | Gemini 1.5 Pro (Sept 2024) | Gemini 1.5 Pro (Sept 2024) | gemini-1-5-pro-sept-2024 | 132.71 | 130.45 | 136.54 | |||
| Qwen2.5-32B | qwen2.5-32b-instruct | Alibaba | China | 2024-09-17 | Open weights (unrestricted) | 2024-09-17 | 3.5e+24 | Confident | Qwen2.5-32B (instruct) | Qwen2.5-32B | qwen2-5-32b | ||||||
| Mistral Large | mistral-large-2402 | Mistral AI | France | 2024-02-26 | API access | 2024-02-26 | 1.1e+25 | Likely | Mistral Large | mistral-large | 120.85 | 114.93 | 128.15 | ||||
| Claude 3 Sonnet | claude-3-sonnet-20240229 | Anthropic | United States of America | 2024-02-29 | API access | 2024-03-04 | Unknown | The mid-sized model in Anthropic's Claude 3 model series. | Claude 3 Sonnet | Claude 3 Sonnet | claude-3-sonnet | 119.69 | 113.60 | 124.18 | |||
| Gemini 1.0 Pro | gemini-1.0-pro-001 | Google DeepMind | United States of America | 2023-12-13 | API access | 2023-12-06 | Speculative | An earlier flagship model from Google DeepMind. | Gemini 1.0 Pro | gemini-1-0-pro | 116.73 | 113.93 | 121.07 | ||||
| Llama 3.1-8B | Llama-3.1-8B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 1.2e+24 | Likely | An instruction-tuned version of the 8 billion-parameter model in Meta’s LLaMA 3.1 series. | Llama 3.1-8B | llama-3-1-8b | 115.08 | 105.54 | 121.34 | |||
| Llama 3.1-405B | Llama-3.1-405B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 3.8e+25 | Confident | An instruction-tuned version of the 405 billion-parameter model in Meta’s LLaMA 3.1 series. | Llama 3.1-405B | llama-3-1-405b | 128.99 | 126.57 | 132.94 | |||
| Qwen2.5-72B | qwen2.5-72b-instruct | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | Confident | An instruction-tuned version of the 72 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-72B | Qwen2.5-72B | qwen2-5-72b | 129.08 | 123.83 | 133.37 | ||
| GPT-4o | gpt-4o-2024-05-13 | GPT-4o (May 2024) | OpenAI | United States of America | 2024-05-13 | API access | 2024-05-13 | Speculative | The first version of GPT-4o, OpenAI's last multimodal model in the GPT-4 series, which powered ChatGPT from mid-2024 to -2025. | GPT-4o (May 2024) | GPT-4o (May 2024) | gpt-4o-may-2024 | 128.78 | 125.93 | 132.66 | ||
| Llama 3.1-70B | Llama-3.1-70B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 7.9e+24 | Confident | An instruction-optimized version of Llama 3.1 70B, the 70 billion-parameter model in Meta's Llama 3.1 series. | Llama 3.1-70B | llama-3-1-70b | 125.42 | 122.66 | 129.56 | |||
| Gemini 1.5 Flash | gemini-1.5-flash-001 | Gemini 1.5 Flash (May 2024) | Google DeepMind | United States of America | 2024-05-23 | API access | 2024-05-10 | Unknown | The first version of Google DeepMind's Gemini 1.5 Flash. | Gemini 1.5 Flash (May 2024) | Gemini 1.5 Flash (May 2024) | gemini-1-5-flash-may-2024 | 122.32 | 118.95 | 125.93 | ||
| Claude 3 Haiku | claude-3-haiku-20240307 | Anthropic | United States of America | 2024-03-07 | API access | 2024-03-04 | Unknown | The smallest model in Anthropic's Claude 3 series. | Claude 3 Haiku | Claude 3 Haiku | claude-3-haiku | 117.25 | 113.54 | 120.58 | |||
| Llama 3-8B | Meta-Llama-3-8B-Instruct | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.2e+23 | Confident | An instruction-tuned version of the 8 billion-parameter model in Meta's Llama 3 series. | Llama 3-8B | llama-3-8b | 116.14 | 112.69 | 118.47 | |||
| Claude 3 Opus | claude-3-opus-20240229 | Anthropic | United States of America | 2024-02-29 | API access | 2024-03-04 | Speculative | The largest model in Anthropic's Claude 3 model series. | Claude 3 Opus | Claude 3 Opus | claude-3-opus | 126.86 | 124.46 | 131.94 | |||
| Phi-4 | phi-4 | Microsoft Research | United States of America | 2024-12-12 | Open weights (unrestricted) | 2024-12-12 | 9.3e+23 | Confident | Phi-4 | Phi-4 | phi-4 | 131.02 | 127.75 | 135.03 | |||
| Mistral Large 2 | mistral-large-2407 | Mistral AI | France | 2024-07-24 | Open weights (non-commercial) | 2024-07-24 | 2.1e+25 | Likely | Mistral Large 2 (Jul 2024) | Mistral Large 2 (Jul 2024) | mistral-large-2-jul-2024 | 127.49 | 124.32 | 131.52 | |||
| phi-3-medium 14B | Phi-3-medium-128k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 4.0e+23 | Likely | phi-3-medium 14B | phi-3-medium-14b | 120.86 | 117.49 | 124.09 | ||||
| GPT-4 Turbo (Nov 2023) | gpt-4-0125-preview | GPT-4 Turbo Preview (January 2024) | OpenAI | United States of America | 2024-01-25 | API access | 2023-11-06 | Unknown | A January 2024 preview version of GPT-4 Turbo, OpenAI's updated GPT-4 series model. | GPT-4 Turbo (Nov 2023) | gpt-4-turbo-nov-2023 | ||||||
| GPT-3.5 Turbo | gpt-3.5-turbo-1106 | GPT-3.5 Turbo (Nov 2023) | OpenAI | United States of America | 2023-11-06 | API access | 2023-06-13 | Likely | A November 2023 preview of GPT-3.5 turbo, OpenAI's updated GPT-3.5 series model. | GPT-3.5 Turbo (Nov 2023) | GPT-3.5 Turbo (Nov 2023) | gpt-3-5-turbo-nov-2023 | 118.00 | 112.63 | 122.23 | ||
| GPT-4 Turbo (Nov 2023) | gpt-4-1106-preview | GPT-4 Turbo Preview (Nov 2023) | OpenAI | United States of America | 2023-11-06 | API access | 2023-11-06 | Unknown | A November 2023 preview of GPT-4 turbo, OpenAI's updated GPT-4 series model. | GPT-4 Turbo (Nov 2023) | gpt-4-turbo-nov-2023 | ||||||
| GPT-3.5 Turbo | gpt-3.5-turbo-0125 | GPT-3.5 Turbo (Jan 2024) | OpenAI | United States of America | 2024-01-25 | API access | 2023-06-13 | Likely | A January 2024 preview version of GPT-3.5 Turbo, OpenAI's updated GPT-3.5 series model. | GPT-3.5 Turbo (Jan 2024) | GPT-3.5 Turbo (Jan 2024) | gpt-3-5-turbo-jan-2024 | 113.60 | 64.33 | 116.41 | ||
| Yi-1.5-34B | Yi-1.5-34B-Chat | 01.AI | China | 2024-05-13 | Open weights (restricted use) | 2024-05-13 | 7.3e+23 | Confident | Yi-1.5-34B (chat) | Yi-1.5-34B | yi-1-5-34b | ||||||
| Yi-34B | Yi-34B-Chat | 01.AI | China | 2023-11-22 | Open weights (restricted use) | 2023-11-02 | 6.1e+23 | Confident | Yi-34B (chat) | Yi-34B | Yi-34B | yi-34b | 116.82 | 111.91 | 120.05 | ||
| Qwen2-72B | qwen2-72b-instruct | Alibaba | China | 2024-06-07 | Open weights (unrestricted) | 2024-06-07 | 3.0e+24 | Confident | Qwen's 72 billion-parameter model in the Qwen 2 series. | Qwen2-72B | Qwen2-72B | qwen2-72b | 125.37 | 122.17 | 128.67 | ||
| Qwen1.5-72B | qwen1.5-72b-chat | Alibaba | China | 2024-02-04 | Open weights (restricted use) | 2024-02-04 | 1.3e+24 | Confident | Qwen1.5-72B | Qwen1.5-72B | qwen1-5-72b | ||||||
| Qwen1.5-32B | qwen1.5-32b-chat | Alibaba | China | 2024-04-03 | Open weights (restricted use) | 2024-02-05 | Confident | A chat-optimized 32 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-32B (chat) | Qwen1.5-32B | qwen1-5-32b | ||||||
| Mistral 7B | Mistral-7B-Instruct-v0.3 | Mistral AI | France | 2024-05-27 | Open weights (unrestricted) | 2023-10-10 | Confident | An instruction-tuned version of Mistral-7B-v0.3, an updated version of their 7 billion-parameter model. | Mistral 7B v0.3 | Mistral 7B v0.3 | mistral-7b-v0-3 | ||||||
| DeepSeek LLM 67B | deepseek-llm-67b-chat | DeepSeek | China | 2023-11-29 | Open weights (restricted use) | 2024-01-05 | 8.0e+23 | Confident | DeepSeek LLM 67B (chat) | DeepSeek LLM 67B | deepseek-llm-67b | ||||||
| Mixtral 8x7B | Mixtral-8x7B-Instruct-v0.1 | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | Speculative | An instruction-tuned version of Mistral's 7 billion active (56 billion total)-parameter mixture-of-experts model | Mixtral 8x7B | mixtral-8x7b | 117.84 | 114.43 | 120.59 | |||
| WizardLM-2 8x22B | WizardLM-2-8x22B | Microsoft | United States of America | 2024-04-15 | Open weights (unrestricted) | 2024-04-15 | Confident | WizardLM-2 8x22B | WizardLM-2 8x22B | wizardlm-2-8x22b | |||||||
| DBRX | dbrx-instruct | Databricks | United States of America | 2024-03-27 | Open weights (restricted use) | 2024-03-27 | 2.6e+24 | Confident | DBRX (instruct) | DBRX | dbrx | ||||||
| Ministral 3B | ministral-3b-2410 | Mistral AI | France | 2024-10-16 | API access | 2024-10-16 | Confident | Ministral 3B | ministral-3b | ||||||||
| Mixtral 8x22B | open-mixtral-8x22b | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | Speculative | Mixtral 8x22B | mixtral-8x22b | 121.08 | 114.24 | 124.41 | ||||
| Ministral 8B | ministral-8b-2410 | Mistral AI | France | 2024-10-16 | Open weights (non-commercial) | 2024-10-16 | Confident | Ministral 8B | ministral-8b | ||||||||
| Mixtral 8x7B | open-mixtral-8x7b | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | Speculative | Mixtral 8x7B | mixtral-8x7b | 117.84 | 114.43 | 120.59 | ||||
| Mistral 7B | open-mistral-7b | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | Confident | Mistral 7B v0.3 | Mistral 7B v0.3 | mistral-7b-v0-3 | |||||||
| Mistral NeMo | open-mistral-nemo-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | Confident | Mistral NeMo | mistral-nemo | 118.13 | 112.83 | 125.53 | |||||
| Eurus-2-7B-PRIME | Eurus-2-7B-PRIME | Tsinghua University,University of Illinois Urbana-Champaign (UIUC),Shanghai AI Lab,Peking University,Shanghai Jiao Tong University,CUHK Shenzhen Research Institute | United States of America,China | 2024-12-31 | Open weights (unrestricted) | 2025-02-03 | Speculative | Eurus-2-7B-PRIME | eurus-2-7b-prime | ||||||||
| Gemini 2.5 Deep Think | gemini-2.5-deep-think-2025-08-01-webapp | Google,Google DeepMind | United States of America | 2025-08-01 | Hosted access (no API) | 2025-08-01 | Unknown | Gemini 2.5 Deep Think | gemini-2-5-deep-think | ||||||||
| Claude Opus 4.7 | claude-opus-4-7 | Claude Opus 4.7 (no thinking) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (no thinking) | Claude Opus 4.7 | claude-opus-4-7 | 156.31 | 153.46 | 163.87 | |||
| Claude Opus 4.7 | claude-opus-4-7_high | Claude Opus 4.7 (high) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (high) | Claude Opus 4.7 | claude-opus-4-7 | 156.31 | 153.46 | 163.87 | |||
| Claude Opus 4.6 | claude-opus-4-6_high | Claude Opus 4.6 (high) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (high) | Claude Opus 4.6 | claude-opus-4-6 | 155.55 | 152.30 | 163.09 | |||
| Claude Opus 4.6 | claude-opus-4-6_unknown | Claude Opus 4.6 (unknown thinking) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (unknown settings) | Claude Opus 4.6 | claude-opus-4-6 | 155.55 | 152.30 | 163.09 | |||
| Claude Opus 4.5 | claude-opus-4-5-20251101_unknown | Claude Opus 4.5 (unknown thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (unknown settings) | Claude Opus 4.5 | claude-opus-4-5 | 149.62 | 146.60 | 155.89 | |||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_unknown | Claude Sonnet 4.5 (unknown thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (59k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 146.85 | 144.23 | 152.58 | |||
| GPT-5.1 | gpt-5.1-2025-11-13_unknown | GPT-5.1 (unknown thinking) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (high) | GPT-5.1 | gpt-5-1 | 149.84 | 147.18 | 156.15 | |||
| Claude Opus 4 | claude-opus-4-20250514_unknown | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 32,000 reasoning tokens. | Claude Opus 4 (unknown settings) | Claude Opus 4 | claude-opus-4 | 143.10 | 140.42 | 148.78 | ||
| Kimi K2 | Kimi-K2-Instruct | Kimi K2 Instruct | Moonshot | China | 2025-07-12 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | Moonshot AI’s 1-trillion parameter scale model. | Kimi K2 (Jul 2025) | Kimi K2 (Jul 2025) | kimi-k2-jul-2025 | 140.42 | 137.12 | 145.88 | |
| Qwen3-Coder-480B-A35B | Qwen3-Coder-480B-A35B-Instruct | Alibaba | China | 2025-07-31 | Open weights (unrestricted) | 2025-07-22 | 1.6e+24 | Confident | Qwen3 Coder | Qwen3-Coder-480B-A35B | qwen3-coder-480b-a35b | ||||||
| Claude Sonnet 4 | claude-sonnet-4-20250514_unknown | Claude Sonnet 4 (unknown thinking) | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized Claude 4 series model, evaluated at up to 12,000 reasoning tokens. | Claude Sonnet 4 (12k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.32 | 139.26 | 147.36 | ||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_unknown | Claude 3.7 Sonnet (unknown thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 64,000 reasoning tokens allowed. | Claude 3.7 Sonnet (64k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 141.83 | 139.34 | 147.11 | |
| GLM-4.5-Air | GLM-4.5-Air | Z.ai (Zhipu AI),Tsinghua University | China | 2025-07-20 | Open weights (unrestricted) | 2025-08-05 | 1.7e+24 | Confident | A smaller version of GLM-4.5, Zhipu AI & Tsinghua University's large open source model. | GLM-4.5-Air | glm-4-5-air | ||||||
| o3-mini | o3-mini-2025-01-31_low | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | Unknown | A smaller version of OpenAI's o3 reasoning model, evaluated with low reasoning effort. | o3-mini | o3-mini | 141.40 | 138.74 | 146.89 | ||||
| video-SALMONN 2+ | video-SALMONN-2plus | ByteDance | China | 2025-06-18 | 2025-06-18 | Likely | video-SALMONN 2+ | video-salmonn-2 | |||||||||
| Gemini 1.5 Pro | gemini-1.5-pro-001-feb24 | Google DeepMind | United States of America | 2024-02-15 | API access | 2024-02-15 | Speculative | The first of two versions of a previous flagship model from Google DeepMind, Gemini 1.5 Pro. | Gemini 1.5 Pro (Feb 2024) | Gemini 1.5 Pro (Feb 2024) | gemini-1-5-pro-feb-2024 | ||||||
| Qwen2.5-72B | Qwen2.5-VL-72B-Instruct | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | Confident | A 72 billion-parameter instruction-tuned vision language model in the Qwen 2.5 series. | Qwen2.5-VL-72B | Qwen2.5-72B | qwen2-5-72b | 129.08 | 123.83 | 133.37 | ||
| InternVL2_5-78B | InternVL2_5-78B | Shanghai AI Lab,SenseTime,Tsinghua University,Nanjing University,Fudan University,Chinese University of Hong Kong (CUHK),Shanghai Jiao Tong University | China,Hong Kong | 2024-12-06 | Open weights (restricted use) | 2024-12-06 | Confident | InternVL2_5-78B | internvl2-5-78b | ||||||||
| Qwen2-VL-72B-Instruct | 2024-08-29 | Unverified | A 72 billion-parameter instruction-tuned vision language model in Qwen's Qwen 2 series. | qwen2-vl-72b-instruct | |||||||||||||
| LLaVA-Video-72B-Qwen2 | 2024-09-02 | Unverified | llava-video-72b-qwen2 | ||||||||||||||
| LinVT | 2024-12-06 | Unverified | linvt | ||||||||||||||
| Aria | 2024-09-30 | Unverified | aria | ||||||||||||||
| ViLAMP-llava-qwen | 2025-05-01 | Unverified | vilamp-llava-qwen | ||||||||||||||
| Oryx-1.5-32B | 2024-10-22 | Unverified | oryx-1-5-32b | ||||||||||||||
| LLaVA-OneVision 72B | 2024-08-06 | Unverified | llava-onevision-72b | ||||||||||||||
| VideoLLAMA3-7B | 2024-01-22 | Unverified | videollama3-7b | ||||||||||||||
| LLaVA-Video-7B-Qwen2 | 2024-09-02 | Unverified | llava-video-7b-qwen2 | ||||||||||||||
| LLaVA-Video-7B-Qwen2-TPO | 2025-01-19 | Unverified | llava-video-7b-qwen2-tpo | ||||||||||||||
| VideoChat-Flash-Qwen2-7B_res448 | 2025-01-11 | Unverified | videochat-flash-qwen2-7b-res448 | ||||||||||||||
| ByteVideoLLM-14B | 2024-10-13 | Unverified | bytevideollm-14b | ||||||||||||||
| NVILA 8B | NVILA-8B | NVIDIA,Massachusetts Institute of Technology (MIT),University of California (UC) Berkeley,University of California San Diego,University of Washington,Tsinghua University | United States of America,China | 2024-12-10 | Open weights (non-commercial) | 2024-12-05 | 2.3e+21 | Likely | NVILA 8B | nvila-8b | |||||||
| LiveCC-7B-Instruct | 2024-04-12 | Unverified | livecc-7b-instruct | ||||||||||||||
| Qwen2-VL-7B-Instruct | 2024-08-29 | Unverified | A 7 billion-parameter instruction-tuned vision language model in Qwen's Qwen 2 series. | qwen2-vl-7b-instruct | |||||||||||||
| MiniCPM-o-2_6 | 2025-01-12 | Unverified | minicpm-o-2-6 | ||||||||||||||
| VideoLLAMA2-7B | 2024-06-12 | Unverified | videollama2-7b | ||||||||||||||
| InternVL2-40B | InternVL2-40B | Shanghai AI Lab | China | 2024-07-08 | Open weights (unrestricted) | 2024-07-04 | Confident | InternVL2-40B | internvl2-40b | ||||||||
| MiniCPM-V-2_6 | 2024-08-03 | Unverified | minicpm-v-2-6 | ||||||||||||||
| GPT-4V | gpt-4-1106-vision-preview | OpenAI | United States of America | 2023-11-06 | API access | 2023-09-25 | Unknown | GPT-4V | gpt-4v | ||||||||
| mPLUG-Owl3-7B-241101 | 2024-11-26 | Unverified | mplug-owl3-7b-241101 | ||||||||||||||
| TimeMarker | 2024-11-27 | Unverified | timemarker | ||||||||||||||
| VITA-1.5 | 2024-12-20 | Unverified | vita-1-5 | ||||||||||||||
| kangaroo | 2024-11-13 | Unverified | kangaroo | ||||||||||||||
| VITA | 2024-08-12 | Unverified | vita | ||||||||||||||
| Video-XL-7B | 2024-10-17 | Unverified | video-xl-7b | ||||||||||||||
| Video-CCAM-7B-v1.2 | 2024-09-29 | Unverified | video-ccam-7b-v1-2 | ||||||||||||||
| long-llava-qwen2-7b | 2024-08-30 | Unverified | long-llava-qwen2-7b | ||||||||||||||
| LongVA-7B | 2024-06-13 | Unverified | longva-7b | ||||||||||||||
| Qwen-VL-Max | 2024-01-18 | Unverified | qwen-vl-max | ||||||||||||||
| InternVL-Chat-V1-5 | 2024-04-18 | Unverified | internvl-chat-v1-5 | ||||||||||||||
| SliME-Llama3-8B | 2024-06-02 | Unverified | slime-llama3-8b | ||||||||||||||
| Chat-Uni-Vi-7B-v1.5 + 100k SG-WV | 2024-06-20 | Unverified | chat-uni-vi-7b-v1-5--100k-sg-wv | ||||||||||||||
| Qwen-VL-Chat | 2023-08-20 | Unverified | qwen-vl-chat | ||||||||||||||
| Chat-UniVi-7B-v1.5 | 2024-04-23 | Unverified | chat-univi-7b-v1-5- | ||||||||||||||
| sharegpt4video-8b | 2024-05-27 | Unverified | sharegpt4video-8b | ||||||||||||||
| Video-LLaVA-7B | 2023-11-17 | Unverified | video-llava-7b | ||||||||||||||
| video_chat2_mistral | 2023-11-29 | Unverified | video-chat2-mistral | ||||||||||||||
| ST-LLM | 2024-03-28 | Unverified | st-llm | ||||||||||||||
| gemini-3-deep-think-preview | 2026-02-12 | Unverified | gemini-3-deep-think-preview | ||||||||||||||
| GPT-5.5 Pro | gpt-5.5-pro_high | GPT-5.5 Pro (high) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 Pro | gpt-5-5-pro | 161.29 | 158.43 | 170.30 | ||||
| GPT-5.5 | gpt-5.5_high | GPT-5.5 (high) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.63 | 156.38 | 165.91 | ||||
| Claude Opus 4.8 | claude-opus-4-8_high | Claude Opus 4.8 | Anthropic | United States of America | 2026-05-28 | API access | 2026-05-28 | Unverified | Claude Opus 4.8 (high) | Claude Opus 4.8 | claude-opus-4-8 | 158.27 | 154.62 | 165.90 | |||
| Claude Opus 4.8 | claude-opus-4-8_medium | Claude Opus 4.8 | Anthropic | United States of America | 2026-05-28 | API access | 2026-05-28 | Unverified | Claude Opus 4.8 (medium) | Claude Opus 4.8 | claude-opus-4-8 | 158.27 | 154.62 | 165.90 | |||
| GPT-5.5 | gpt-5.5_medium | GPT-5.5 (medium) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.63 | 156.38 | 165.91 | ||||
| Claude Opus 4.7 | claude-opus-4-7_medium | Claude Opus 4.7 (medium) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (medium) | Claude Opus 4.7 | claude-opus-4-7 | 156.31 | 153.46 | 163.87 | |||
| Grok 4.20 | grok-4-20 | xAI | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Grok 4.20 | grok-4-20 | 154.25 | 148.83 | 161.27 | |||||
| Claude Opus 4.8 | claude-opus-4-8_low | Claude Opus 4.8 | Anthropic | United States of America | 2026-05-28 | API access | 2026-05-28 | Unverified | Claude Opus 4.8 (low) | Claude Opus 4.8 | claude-opus-4-8 | 158.27 | 154.62 | 165.90 | |||
| Claude Opus 4.7 | claude-opus-4-7_low | Claude Opus 4.7 (low) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (low) | Claude Opus 4.7 | claude-opus-4-7 | 156.31 | 153.46 | 163.87 | |||
| Claude Sonnet 4.6 | claude-sonnet-4-6_high | Claude Sonnet 4.6 (high) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 153.41 | 150.31 | 160.16 | ||||
| Claude Sonnet 4.6 | claude-sonnet-4-6_max | Claude Sonnet 4.6 (max) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 153.41 | 150.31 | 160.16 | ||||
| GPT-5.4 | gpt-5.4-2026-03-05_medium | GPT-5.4 (medium) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (medium) | GPT-5.4 | gpt-5-4 | 156.22 | 153.71 | 163.79 | |||
| GPT-5.2 Pro | gpt-5.2-pro-2025-12-11_high | GPT-5.2 Pro | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Likely | GPT-5.2 Pro (high) | GPT-5.2 Pro | gpt-5-2-pro | 154.72 | 152.20 | 161.41 | |||
| GPT-5.2 Pro | gpt-5.2-pro-2025-12-11_medium | GPT-5.2 Pro | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Likely | GPT-5.2 Pro (medium) | GPT-5.2 Pro | gpt-5-2-pro | 154.72 | 152.20 | 161.41 | |||
| Claude Opus 4.5 | claude-opus-4-5-20251101_64K | Claude Opus 4.5 (64k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (64k thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.62 | 146.60 | 155.89 | |||
| GPT-5.4 | gpt-5.4-2026-03-05_low | GPT-5.4 (low) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (low) | GPT-5.4 | gpt-5-4 | 156.22 | 153.71 | 163.79 | |||
| GPT-5 Pro | gpt-5-pro-2025-10-06_unknown | GPT-5 Pro | OpenAI | United States of America | 2025-10-07 | API access | 2025-10-07 | Unknown | GPT-5 Pro (unknown thinking) | GPT-5 Pro | gpt-5-pro | 150.48 | 147.41 | 157.34 | |||
| Claude Opus 4.5 | claude-opus-4-5-20251101_8K | Claude Opus 4.5 (8k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (8k thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.62 | 146.60 | 155.89 | |||
| Gemini 3.5 Flash | gemini-3.5-flash_minimal | Gemini 3.5 Flash (minimal) | United States of America | 2026-05-19 | Unverified | Gemini 3.5 Flash | gemini-3-5-flash | 154.89 | 152.16 | 162.79 | |||||||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_8K | Claude Sonnet 4.5 (8k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (2k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 146.85 | 144.23 | 152.58 | |||
| tiny-recursion-model | 2025-10-06 | Unverified | tiny-recursion-model | ||||||||||||||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_1K | Claude Sonnet 4.5 (1k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (2k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 146.85 | 144.23 | 152.58 | |||
| GPT-5.4 Nano | gpt-5.4-nano-2026-03-17_xhigh | GPT-5.4 nano (xhigh) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 nano (xhigh) | GPT-5.4 Nano | gpt-5-4-nano | 146.54 | 143.59 | 152.29 | |||
| Grok 4 Fast | grok-4-fast | xAI | United States of America | 2025-09-19 | API access | 2025-09-19 | Unknown | Grok 4 Fast | Grok 4 Fast | grok-4-fast | 144.72 | 141.85 | 151.94 | ||||
| o3-pro | o3-pro-2025-06-10_high | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | Unknown | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with high reasoning effort. | o3-pro | o3-pro | 148.04 | 145.41 | 155.24 | |||||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro_32K | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-05 | Unknown | Google’s general-purpose frontier reasoning model, evaluated with up to 32,000 reasoning tokens. | Gemini 2.5 Pro (32k thinking) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.17 | 143.36 | 151.99 | ||
| MiniMax-M2.5 | MiniMax-M2.5 | MiniMax | China | 2026-02-12 | Open weights (unrestricted) | 2026-02-12 | Likely | MiniMax-M2.5 | minimax-m2-5 | 147.22 | 141.58 | 152.64 | |||||
| Claude Opus 4 | claude-opus-4-20250514_8K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 8,000 reasoning tokens. | Claude Opus 4 (8k thinking) | Claude Opus 4 | claude-opus-4 | 143.10 | 140.42 | 148.78 | ||
| GPT-5.4 Mini | gpt-5.4-mini-2026-03-17_medium | GPT-5.4 mini (medium) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 mini (medium) | GPT-5.4 Mini | gpt-5-4-mini | 148.96 | 145.80 | 155.00 | |||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro_16K | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-05 | Unknown | Google’s general-purpose frontier reasoning model, evaluated with up to 16,000 reasoning tokens. | Gemini 2.5 Pro (16k thinking) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.17 | 143.36 | 151.99 | ||
| DeepSeek-V3.2 | deepseek/deepseek-v3.2 | DeepSeek-V3.2 (Thinking; Novita) | DeepSeek | China | 2025-12-01 | Open weights (unrestricted) | 2025-12-01 | 4.2e+24 | Likely | DeepSeek-V3.2 | DeepSeek-V3.2 | deepseek-v3-2 | 146.64 | 143.60 | 152.74 | ||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro_8K | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-05 | Unknown | Google’s general-purpose frontier reasoning model, evaluated with up to 8,000 reasoning tokens. | Gemini 2.5 Pro (8k thinking) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.17 | 143.36 | 151.99 | ||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_16K | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (16k thinking) | Claude Haiku 4.5 | claude-haiku-4-5 | 142.91 | 139.17 | 148.79 | ||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_23K | Google DeepMind | United States of America | 2025-05-20 | API access | 2025-04-17 | Unknown | A May 2025 preview version of Gemini 2.5 Flash, a small model in Google DeepMind's Gemini 2.5 series, evaluated with up to 23,000 reasoning tokens. | Gemini 2.5 Flash (May 2025) | Gemini 2.5 Flash (May 2025) | gemini-2-5-flash-may-2025 | 142.30 | 139.70 | 147.65 | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_1K | Google DeepMind | United States of America | 2025-05-20 | API access | 2025-04-17 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 1,000 reasoning tokens. | Gemini 2.5 Flash (May 2025) | Gemini 2.5 Flash (May 2025) | gemini-2-5-flash-may-2025 | 142.30 | 139.70 | 147.65 | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_8K | Google DeepMind | United States of America | 2025-05-20 | API access | 2025-04-17 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 8,000 reasoning tokens. | Gemini 2.5 Flash (May 2025) | Gemini 2.5 Flash (May 2025) | gemini-2-5-flash-may-2025 | 142.30 | 139.70 | 147.65 | |||
| Claude Sonnet 4 | claude-sonnet-4-20250514_8K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Claude Sonnet 4 (8k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.32 | 139.26 | 147.36 | ||||
| o3-pro | o3-pro-2025-06-10_low | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | Unknown | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with low reasoning effort. | o3-pro | o3-pro | 148.04 | 145.41 | 155.24 | |||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_16K | Google DeepMind | United States of America | 2025-05-20 | API access | 2025-04-17 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 16,000 reasoning tokens. | Gemini 2.5 Flash (May 2025) | Gemini 2.5 Flash (May 2025) | gemini-2-5-flash-may-2025 | 142.30 | 139.70 | 147.65 | |||
| o3-pro | o3-pro-2025-06-10_medium | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | Unknown | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with medium reasoning effort. | o3-pro | o3-pro | 148.04 | 145.41 | 155.24 | |||||
| GPT-5 | gpt-5-2025-08-07_low | GPT-5 (low) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using low reasoning effort. | GPT-5 (low) | GPT-5 | gpt-5 | 150.00 | 147.33 | 156.34 | |
| GPT-5.4 Nano | gpt-5.4-nano-2026-03-17_medium | GPT-5.4 nano (medium) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 nano (medium) | GPT-5.4 Nano | gpt-5-4-nano | 146.54 | 143.59 | 152.29 | |||
| GPT-5 mini | gpt-5-mini-2025-08-07_minimal | GPT-5 mini (minimal) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 mini (medium) | GPT-5 mini | gpt-5-mini | 145.83 | 143.09 | 151.48 | ||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_8K | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (8k thinking) | Claude Haiku 4.5 | claude-haiku-4-5 | 142.91 | 139.17 | 148.79 | ||||
| GPT-5.4 Nano | gpt-5.4-nano-2026-03-17_low | GPT-5.4 nano (low) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 nano (low) | GPT-5.4 Nano | gpt-5-4-nano | 146.54 | 143.59 | 152.29 | |||
| codex-mini-2025-05-16 | 2025-05-16 | Unverified | codex-mini-2025-05-16 | ||||||||||||||
| Qwen3-235B-A22B (Jul 2025) | Qwen3-235B-A22B-Instruct-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | A 22 billion active (235 billion total)-parameter mixture-of-experts instruction-tuned model in the Qwen 3 series. | Qwen3 Non-thinking (Jul 2025) | Qwen3-235B-A22B-Instruct (Jul 2025) | Qwen3-235B-A22B-Instruct (Jul 2025) | qwen3-235b-a22b-instruct-jul-2025 | 138.99 | 135.96 | 144.81 | |
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_1K | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (1k thinking) | Claude Haiku 4.5 | claude-haiku-4-5 | 142.91 | 139.17 | 148.79 | ||||
| GPT-5.4 Mini | gpt-5.4-mini-2026-03-17_low | GPT-5.4 mini (low) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 mini (low) | GPT-5.4 Mini | gpt-5-4-mini | 148.96 | 145.80 | 155.00 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_8K | Claude 3.7 Sonnet (8k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 9,000 reasoning tokens allowed. | Claude 3.7 Sonnet (8k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 141.83 | 139.34 | 147.11 | |
| Claude Sonnet 4 | claude-sonnet-4-20250514_1K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized Claude 4 series model, evaluated at up to 1,000 reasoning tokens. | Claude Sonnet 4 (1k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.32 | 139.26 | 147.36 | |||
| o1-mini | o1-mini-2024-09-12_unknown | o1-mini (unknown thinking) | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | Unknown | A smaller version of OpenAI’s first reasoning model, o1, evaluated with high reasoning effort. | o1-mini | o1-mini | 136.56 | 133.72 | 141.20 | |||
| GPT-5.2 | gpt-5.2-2025-12-11_unknown | GPT-5.2 (unknown thinking) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (high) | GPT-5.2 | gpt-5-2 | 153.75 | 151.23 | 160.30 | |||
| GPT-5 mini | gpt-5-mini-2025-08-07_low | GPT-5 mini (low) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 mini (medium) | GPT-5 mini | gpt-5-mini | 145.83 | 143.09 | 151.48 | ||
| Grok-3 mini | grok-3-mini_low | Grok 3 mini | xAI | United States of America | 2025-06-24 | API access | 2025-02-19 | Unknown | Grok 3 mini (low) | Grok-3 mini | grok-3-mini | 141.05 | 138.50 | 145.60 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_1K | Claude 3.7 Sonnet (1k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 1,000 reasoning tokens allowed. | Claude 3.7 Sonnet (1k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 141.83 | 139.34 | 147.11 | |
| Llama 4 Maverick | Llama-4-Maverick-17B-128E-Instruct | Llama 4 Maverick | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | Likely | The 17 billion active (400 billion total)-parameter model in Meta's Llama 4 series. | Llama 4 Maverick | llama-4-maverick | 133.04 | 130.31 | 137.25 | ||
| Grok 3 | grok-3 | xAI | United States of America | 2025-04-09 | API access | 2025-02-17 | 3.5e+26 | Likely | XAI's third generation flagship model. | Grok 3 | Grok 3 | grok-3 | 139.06 | 136.97 | 143.60 | ||
| Claude Opus 4 | claude-opus-4-20250514_1K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 1,000 reasoning tokens. | Claude Opus 4 (1k thinking) | Claude Opus 4 | claude-opus-4 | 143.10 | 140.42 | 148.78 | ||
| Magistral Medium 1.1 | magistral-medium-2506 | Mistral AI | France | 2025-06-10 | API access | 2025-06-10 | Unknown | Magistral Medium 1.1 | magistral-medium-1-1 | ||||||||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro-preview-06-05_1K | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America | 2025-06-05 | API access | 2025-06-05 | Unknown | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model, evaluated with up to 1,000 reasoning tokens. | Gemini 2.5 Pro Preview (Jun 2025, 1k thinking) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.17 | 143.36 | 151.99 | |
| GPT-5 | gpt-5-2025-08-07_minimal | GPT-5 (minimal) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using minimal reasoning effort. | GPT-5 (minimal) | GPT-5 | gpt-5 | 150.00 | 147.33 | 156.34 | |
| GPT-5 nano | gpt-5-nano-2025-08-07_low | GPT-5 nano (low) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 nano (low) | GPT-5 nano | gpt-5-nano | 140.85 | 136.04 | 146.03 | ||
| GPT-5 nano | gpt-5-nano-2025-08-07_minimal | GPT-5 nano (minimal) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 nano (low) | GPT-5 nano | gpt-5-nano | 140.85 | 136.04 | 146.03 | ||
| Claude Fable 5 | claude-fable-5_high | Claude Fable 5 (high) | Anthropic | United States of America | 2026-06-09 | API access | 2026-06-09 | Confident | Claude Fable 5 (high) | Claude Fable 5 | claude-fable-5 | 161.00 | 158.16 | 169.22 | |||
| Claude Opus 4.8 | claude-opus-4-8_xhigh | Claude Opus 4.8 | Anthropic | United States of America | 2026-05-28 | API access | 2026-05-28 | Unverified | Claude Opus 4.8 (xhigh) | Claude Opus 4.8 | claude-opus-4-8 | 158.27 | 154.62 | 165.90 | |||
| GLM-5.2 | glm-5.2_high | GLM-5.2 | Z.ai (Zhipu AI) | China | 2026-06-16 | Open weights (unrestricted) | 2026-06-16 | Unverified | GLM-5.2 (high) | GLM-5.2 | glm-5-2 | 152.72 | 149.02 | 159.70 | |||
| Gemini 3.5 Flash | gemini-3.5-flash_unknown | Gemini 3.5 Flash (unknown thinking) | United States of America | 2026-05-19 | Unverified | Gemini 3.5 Flash | gemini-3-5-flash | 154.89 | 152.16 | 162.79 | |||||||
| Claude Sonnet 4.6 | claude-sonnet-4-6_medium | Claude Sonnet 4.6 (medium) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 153.41 | 150.31 | 160.16 | ||||
| Claude Fable 5 | claude-fable-5_unknown | Claude Fable 5 (unknown) | Anthropic | United States of America | 2026-06-09 | API access | 2026-06-09 | Confident | Claude Fable 5 (unknown) | Claude Fable 5 | claude-fable-5 | 161.00 | 158.16 | 169.22 | |||
| Grok 4.3 Beta | grok-4-3 | xAI | United States of America | 2026-04-17 | Hosted access (no API) | 2026-04-17 | Likely | Grok 4.3 Beta | grok-4-3-beta | 148.94 | 146.24 | 155.28 | |||||
| QwQ-32B | QwQ-32B | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B | QwQ-32B | QwQ-32B | qwq-32b | |||||
| Gemini 2.0 Flash | gemini-exp-1206 | Google DeepMind,Google | United States of America | 2024-12-06 | API access | 2024-12-11 | Unknown | A December 2024 experimental version of Gemini 2.0 Flash, a small langauge model from Google DeepMind. | Gemini 2.0 Flash (Dec 2024) | Gemini 2.0 Flash (Dec 2024) | gemini-2-0-flash-dec-2024 | 135.16 | 123.66 | 139.37 | |||
| Qwen2.5-Max | qwen2.5-max | Alibaba | China | 2025-01-28 | API access | 2025-01-28 | Unknown | The largest model in the Qwen 2.5 series. | Qwen2.5-Max | Qwen2.5-Max | qwen2-5-max | 133.41 | 130.61 | 139.10 | |||
| Gemini 2.0 Flash | gemini-2.0-flash-exp | Google DeepMind,Google | United States of America | 2024-12-11 | API access | 2024-12-11 | Unknown | A December 2024 experimental version of Gemini 2.0 Flash, a small langauge model from Google DeepMind. | Gemini 2.0 Flash (Dec 2024) | Gemini 2.0 Flash (Dec 2024) | gemini-2-0-flash-dec-2024 | 135.16 | 123.66 | 139.37 | |||
| Gemini 2.0 Flash-Lite | gemini-2.0-flash-lite | Google DeepMind | United States of America | 2025-02-05 | API access | 2024-02-05 | Unknown | The smallest model in Google DeepMind's Gemini 2.0 series. | Gemini 2.0 Flash-Lite | gemini-2-0-flash-lite | |||||||
| Gemini 2.0 Flash-Lite | gemini-2.0-flash-lite-preview-02-05 | Google DeepMind | United States of America | 2025-02-05 | API access | 2024-02-05 | Unknown | A February 2025 preview version of the smallest model in Google DeepMind's Gemini 2.0 series. | Gemini 2.0 Flash-Lite | gemini-2-0-flash-lite | |||||||
| Dracarys2-72B-Instruct | 2024-09-30 | Unverified | dracarys2-72b-instruct | ||||||||||||||
| learnlm-1.5-pro-experimental | Unverified | learnlm-1-5-pro-experimental | |||||||||||||||
| sonar | Perplexity Sonar | Unverified | sonar | ||||||||||||||
| Qwen2.5-Coder (32B) | Qwen2.5-Coder-32B-Instruct | Alibaba | China | 2024-11-21 | Open weights (unrestricted) | 2024-11-12 | 1.1e+24 | Confident | An instruction-tuned and coding-optimized 32 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-Coder-32B (instruct) | Qwen2.5-Coder-32B-Instruct | Qwen2.5-Coder-32B-Instruct | qwen2-5-coder-32b-instruct | ||||
| Dracarys2-Llama-3.1-70B-Instruct | 2024-08-14 | Unverified | dracarys2-llama-3-1-70b-instruct | ||||||||||||||
| DeepSeek-R1-Distill-Qwen-32B | DeepSeek-R1-Distill-Qwen-32B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | Speculative | A model based on Qwen 32B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | DeepSeek-R1-Distill-Qwen-32B | deepseek-r1-distill-qwen-32b | |||||||
| Amazon Nova Pro | amazon.nova-pro-v1:0 | Amazon | United States of America | 2024-12-03 | API access | 2024-12-03 | 6.0e+24 | Speculative | Amazon Nova Pro | amazon-nova-pro | 123.73 | 112.50 | 127.72 | ||||
| QwQ-32B | QwQ-32B-Preview | Alibaba | China | 2024-11-28 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B Preview | QwQ-32B-Preview | QwQ-32B-Preview | qwq-32b-preview | |||||
| Amazon Nova Lite | amazon.nova-lite-v1:0 | Amazon | United States of America | 2024-12-03 | API access | 2024-12-03 | Unknown | Amazon Nova Lite | amazon-nova-lite | ||||||||
| Command R+ | c4ai-command-r-plus-08-2024 | Cohere,Cohere for AI | Canada | 2024-08-30 | Open weights (non-commercial) | 2024-04-04 | Confident | Command R+ | Command R+ | command-r | |||||||
| Amazon Nova Micro | amazon.nova-micro-v1:0 | Amazon | United States of America | 2024-12-03 | API access | 2024-12-03 | Unknown | Amazon Nova Micro | amazon-nova-micro | ||||||||
| c4ai-command-r-08-2024 | 2024-08-30 | Unverified | c4ai-command-r-08-2024 | ||||||||||||||
| phi-3-small 7.4B | Phi-3-small-8k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 2.1e+23 | Confident | A 7 billion-parameter model in Microsoft's Phi-3 open source model series with a 8000-token context length. | phi-3-small 7.4B | phi-3-small-7-4b | 121.43 | 117.91 | 124.20 | |||
| phi-3-mini 3.8B | Phi-3-mini-4k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 7.5e+22 | Confident | A 4 billion-parameter model in Microsoft's Phi-3 open source model series with a 4000-token context length. | phi-3-mini 3.8B | phi-3-mini-3-8b | 116.79 | 112.27 | 121.42 | |||
| OLMo 2 Furious 13B | OLMo-2-1124-13B-Instruct | Allen Institute for AI,University of Washington,New York University (NYU) | United States of America | 2024-12-31 | Open weights (unrestricted) | 2024-12-31 | 4.6e+23 | Confident | OLMo 2 Furious 13B | olmo-2-furious-13b | |||||||
| Claude Opus 4.7 | claude-opus-4-7_unknown | Claude Opus 4.7 (unknown) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (unknown settings) | Claude Opus 4.7 | claude-opus-4-7 | 156.31 | 153.46 | 163.87 | |||
| Claude Opus 4.8 | claude-opus-4-8_unknown | Claude Opus 4.8 | Anthropic | United States of America | 2026-05-28 | API access | 2026-05-28 | Unverified | Claude Opus 4.8 (unknown thinking) | Claude Opus 4.8 | claude-opus-4-8 | 158.27 | 154.62 | 165.90 | |||
| Claude Opus 4.8 | claude-opus-4-8_none | Claude Opus 4.8 | Anthropic | United States of America | 2026-05-28 | API access | 2026-05-28 | Unverified | Claude Opus 4.8 (no thinking) | Claude Opus 4.8 | claude-opus-4-8 | 158.27 | 154.62 | 165.90 | |||
| MiniMax-M3 | MiniMax-M3 | MiniMax | China | 2026-06-01 | API access | 2026-06-01 | Confident | MiniMax-M3 | minimax-m3 | ||||||||
| MiMo-V2.5-Pro | mimo-v2.5-pro | Xiaomi Corp | China | Open weights (unrestricted) | 2026-04-23 | 6.8e+24 | Likely | MiMo-V2.5-Pro | mimo-v2-5-pro | ||||||||
| DeepSeek-V4-Pro | deepseek-v4-pro_unknown | DeepSeek v4 Pro (unknown thinking) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 9.7e+24 | Likely | DeepSeek v4 (unknown thinking) | DeepSeek-V4-Pro | deepseek-v4-pro | 149.73 | 146.94 | 156.88 | ||
| GPT-5.5 | gpt-5.5_unknown | GPT-5.5 (unknown thinking) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.63 | 156.38 | 165.91 | ||||
| MiMo-V2-Pro | mimo-v2-pro | Xiaomi Corp | China | 2026-03-18 | Confident | MiMo-V2-Pro | mimo-v2-pro | ||||||||||
| MiMo-V2.5 | mimo-v2.5 | Xiaomi Corp | China | Open weights (unrestricted) | 2026-04-23 | 4.3e+24 | Unverified | MiMo-V2.5 | mimo-v2-5 | ||||||||
| Kimi K2.5 | kimi-k2.5_none | Kimi K2.5 (instant) | Moonshot | China | 2026-01-27 | Open weights (unrestricted) | 2026-02-02 | 5.8e+24 | Likely | Kimi K2.5 | kimi-k2-5 | 148.38 | 145.46 | 154.65 | |||
| GPT-5.3 Codex | gpt-5.3-codex | GPT-5.3 Codex | OpenAI | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | GPT-5.3 Codex | GPT-5.3 Codex | gpt-5-3-codex | 155.85 | 153.78 | 164.12 | |||
| MiniMax-M2.7 | MiniMax-M2.7 | MiniMax | China | 2026-03-18 | Open weights (non-commercial) | 2026-03-18 | Unverified | MiniMax-M2.7 | minimax-m2-7 | 145.82 | 137.48 | 152.17 | |||||
| MiniMax-M2.1 | MiniMax-M2.1 | MiniMax | China | 2025-12-23 | Open weights (restricted use) | 2025-12-23 | Confident | MiniMax-M2.1 | minimax-m2-1 | ||||||||
| GPT-5.4 | gpt-5.4-2026-03-05_unknown | GPT-5.4 | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (none) | GPT-5.4 | gpt-5-4 | 156.22 | 153.71 | 163.79 | |||
| Claude Opus 4.1 | claude-opus-4-1-20250805_unknown | Claude Opus 4.1 (unknown thinking) | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model, benchmarked with up to 27,000 reasoning tokens. | Claude Opus 4.1 (unknown settings) | Claude Opus 4.1 | claude-opus-4-1 | 144.34 | 141.71 | 150.46 | ||
| DeepSeek-V3.2-Exp | DeepSeek-V3.2-Exp_thinking | DeepSeek | China | 2025-09-29 | Open weights (unrestricted) | 2025-09-29 | 4.2e+24 | Confident | DeepSeek-V3.2-Exp | deepseek-v3-2-exp | 144.94 | 142.23 | 150.68 | ||||
| GLM-4.6 | glm-4.6 | GLM-4.6 | Z.ai (Zhipu AI),Tsinghua University | China | 2025-09-30 | Open weights (unrestricted) | 2025-09-30 | 4.4e+24 | Likely | GLM-4.6 | glm-4-6 | 140.96 | 132.50 | 147.47 | |||
| Gemma 4 26B A4B | gemma-4-26b-a4b | Gemma 4 26B A4B | Google DeepMind | United States of America | 2026-04-02 | Open weights (unrestricted) | 2026-04-02 | Confident | Gemma 4 26B A4B | gemma-4-26b-a4b | |||||||
| Qwen3.5-27B | qwen3.5-27B | Alibaba | China | 2026-02-24 | Open weights (restricted use) | 2026-02-24 | Confident | Qwen3.5-27B | qwen3-5-27b | ||||||||
| MiMo-V2-Flash | mimo-v2-flash | Xiaomi Corp | China | Open weights (unrestricted) | 2025-12-16 | 2.4e+24 | Likely | MiMo-V2-Flash | mimo-v2-flash | ||||||||
| GPT-5.2 Codex | gpt-5.2-codex | GPT-5.2 Codex | OpenAI | United States of America | 2025-12-18 | API access | 2025-12-18 | Likely | GPT-5.2 Codex | GPT-5.2 Codex | gpt-5-2-codex | ||||||
| GPT-5.1-Codex | gpt-5.1-codex | GPT-5.1 Codex | OpenAI | United States of America | 2025-11-12 | API access | 2025-11-12 | Unknown | GPT-5.1-codex | GPT-5.1-Codex | gpt-5-1-codex | ||||||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_unknown | Claude Haiku 4.5 (unknown thinking) | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (16k thinking) | Claude Haiku 4.5 | claude-haiku-4-5 | 142.91 | 139.17 | 148.79 | |||
| MiniMax-M2 | MiniMax-M2 | MiniMax | China | 2025-10-27 | Open weights (unrestricted) | 2025-10-27 | Confident | MiniMax-M2 | minimax-m2 | ||||||||
| DeepSeek-V3.2-Exp | DeepSeek-V3.2-Exp | DeepSeek | China | 2025-09-29 | Open weights (unrestricted) | 2025-09-29 | 4.2e+24 | Confident | DeepSeek-V3.2-Exp | deepseek-v3-2-exp | 144.94 | 142.23 | 150.68 | ||||
| Qwen 3.6 35B-A3B | qwen3.6-35b-a3b | Alibaba | China | 2026-04-14 | Open weights (unrestricted) | 2026-04-14 | Unverified | Qwen 3.6 35B-A3B | qwen-3-6-35b-a3b | ||||||||
| Gemini 3.1 Flash-Lite | gemini-3.1-flash-lite | United States of America | 2026-03-03 | API access | 2026-03-03 | Likely | Gemini 3.1 Flash-Lite | gemini-3-1-flash-lite | 144.66 | 104.49 | 149.77 | ||||||
| GPT-5.1-Codex-mini | gpt-5.1-codex-mini | GPT-5.1 Codex Mini | OpenAI | United States of America | 2025-11-12 | API access | 2025-11-12 | Unknown | GPT-5.1-codex mini | GPT-5.1-Codex-mini | gpt-5-1-codex-mini | ||||||
| Grok 4.1 Fast | grok-4-1-fast-reasoning | xAI | United States of America | 2025-11-19 | API access | 2025-11-19 | Unknown | Grok 4.1 Fast | grok-4-1-fast | ||||||||
| Grok 4.1 | grok-4-1 | xAI | United States of America | 2025-11-17 | API access | 2025-11-17 | Unknown | Grok 4.1 | grok-4-1 | ||||||||
| Mercury 2 | mercury-2 | Mercury 2 | Inception Labs | United States of America | API access | 2026-02-20 | Unverified | Mercury 2 | mercury-2 | ||||||||
| grok-code-fast-1 | 2025-08-28 | Unverified | grok-code-fast-1 | ||||||||||||||
| GPT-5.3 Codex | gpt-5.3-codex_xhigh | GPT-5.3 Codex (xhigh) | OpenAI | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | GPT-5.3 Codex | gpt-5-3-codex | 155.85 | 153.78 | 164.12 | ||||
| GPT-5.1-Codex-Max | gpt-5.1-codex-max | GPT-5.1-Codex-Max | OpenAI | United States of America | 2025-11-19 | API access | 2025-11-19 | Unknown | GPT-5.1-Codex-Max | GPT-5.1-Codex-Max | gpt-5-1-codex-max | ||||||
| GPT-5.5 | gpt-5.5_none | GPT-5.5 (no thinking) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.63 | 156.38 | 165.91 | ||||
| GPT-5.4 | gpt-5.4-2026-03-05_none | GPT-5.4 (none) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (none) | GPT-5.4 | gpt-5-4 | 156.22 | 153.71 | 163.79 | |||
| DeepSeek-V4-Pro | deepseek-v4-pro_high | DeepSeek v4 Pro (high) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 9.7e+24 | Likely | DeepSeek v4 (high) | DeepSeek-V4-Pro | deepseek-v4-pro | 149.73 | 146.94 | 156.88 | ||
| DeepSeek-V4-Flash | deepseek-v4-flash_high | DeepSeek v4 Flash (high) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 2.5e+24 | Likely | DeepSeek v4 (high) | DeepSeek-V4-Flash | deepseek-v4-flash | |||||
| Kimi K2 Thinking | kimi-k2-thinking | Kimi K2 Thinking | Moonshot | China | 2025-11-06 | Open weights (restricted use) | 2025-11-06 | 4.2e+24 | Likely | Kimi K2 Thinking | kimi-k2-thinking | 146.08 | 142.98 | 151.63 | |||
| gpt-oss-120b | gpt-oss-120b | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | The larger one of OpenAI's two latest open source language models. | gpt-oss-120b | gpt-oss-120b | 140.55 | 135.47 | 146.61 | |||
| gpt-oss-20b | gpt-oss-20b | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | Confident | The smaller one of OpenAI's two latest open source language models. | gpt-oss-20b | gpt-oss-20b | ||||||
| DeepSeek-V4-Pro | deepseek-v4-pro_none | DeepSeek v4 Pro (unknown thinking) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 9.7e+24 | Likely | DeepSeek v4 (no thinking) | DeepSeek-V4-Pro | deepseek-v4-pro | 149.73 | 146.94 | 156.88 | ||
| qwen3-coder-plus | 2025-09-23 | Unverified | qwen3-coder-plus | ||||||||||||||
| DeepSeek-V4-Flash | deepseek-v4-flash_none | DeepSeek v4 Flash (high) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 2.5e+24 | Likely | DeepSeek v4 (no thinking) | DeepSeek-V4-Flash | deepseek-v4-flash | |||||
| Kimi K2 | moonshotai/kimi-k2-0905 | Kimi K2 0905 (Novita) | Moonshot | China | 2025-09-05 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | A July 2025 preview version of Kimi K2 turbo, an updated version of Kimi K2. | Kimi K2 (Sep 2025) | Kimi K2 (Sep 2025) | kimi-k2-sep-2025 | 141.04 | 137.75 | 147.36 | |
| GPT-5.2 | gpt-5.2-2025-12-11_none | GPT-5.2 (none) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (none) | GPT-5.2 | gpt-5-2 | 153.75 | 151.23 | 160.30 | |||
| Grok 4 | grok-4-0709_high | xAI | United States of America | 2025-07-09 | API access | 2025-07-09 | 5.0e+26 | Speculative | Grok 4 | Grok 4 | grok-4 | 147.10 | 144.69 | 153.68 | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-09-2025 | Google DeepMind | United States of America | 2025-09-25 | API access | 2025-04-17 | Unknown | Gemini 2.5 Flash (Sep 2025) | Gemini 2.5 Flash (Sep 2025) | Gemini 2.5 Flash (Sep 2025) | gemini-2-5-flash-sep-2025 | 143.18 | 134.59 | 150.09 | |||
| MPT-7B | mpt-7b | MosaicML | United States of America | 2023-05-05 | Open weights (unrestricted) | 2023-05-05 | 4.2e+22 | Confident | MPT-7B | mpt-7b | 93.21 | 88.55 | 99.04 | ||||
| Falcon-7B | falcon-7b | Technology Innovation Institute | United Arab Emirates | 2023-04-24 | Open weights (unrestricted) | 2023-04-24 | 6.3e+22 | Confident | The 7 billion-parameter model in the Falcon series. | Falcon-7B | falcon-7b | 93.74 | 88.35 | 99.78 | |||
| chatglm2-6b | 2023-06-24 | Unverified | chatglm2-6b | ||||||||||||||
| internlm-7b | 2023-07-05 | Unverified | internlm-7b | ||||||||||||||
| internlm-20b | 2023-09-18 | Unverified | internlm-20b | ||||||||||||||
| Baichuan 2-7B | Baichuan-2-7B-Base | Baichuan | China | 2023-09-20 | Open weights (restricted use) | 2023-09-20 | 1.1e+23 | Confident | Baichuan 2-7B | baichuan-2-7b | 95.01 | 86.71 | 101.71 | ||||
| Baichuan2-13B | Baichuan-2-13B-Base | Baichuan | China | 2023-09-06 | Open weights (restricted use) | 2023-09-06 | 2.0e+23 | Confident | Baichuan2-13B | baichuan2-13b | 102.12 | 94.55 | 107.23 | ||||
| LLaMA-7B | LLaMA-7B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-24 | 4.0e+22 | Confident | Meta's smallest model in the original Llama series. | LLaMA-7B | llama-7b | 95.39 | 91.08 | 99.17 | |||
| LLaMA-13B | LLaMA-13B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-27 | 7.8e+22 | Confident | Meta's 13 billion-parameter model in the original Llama series. | LLaMA-13B | llama-13b | 99.41 | 95.78 | 102.96 | |||
| LLaMA-33B | LLaMA-33B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-27 | 2.7e+23 | Confident | Meta's 33 billion-parameter model in the original Llama series | LLaMA-33B | llama-33b | 106.49 | 102.78 | 109.15 | |||
| LLaMA-65B | LLaMA-65B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-24 | 5.5e+23 | Confident | Meta's largest model in the original Llama series. | LLaMA-65B | llama-65b | 109.35 | 105.76 | 111.66 | |||
| Llama 2-7B | Llama-2-7b | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.4e+22 | Confident | Meta's 7 billion-parameter model in the Llama 2 series of open source models. | Llama 2-7B | llama-2-7b | 97.84 | 92.86 | 101.65 | |||
| Llama 2-13B | Llama-2-13b | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 1.6e+23 | Confident | Meta's 13 billion-parameter model in the Llama 2 series of open source models. | Llama 2-13B | llama-2-13b | 105.20 | 102.56 | 108.98 | |||
| Llama 2-70B | Llama-2-70b-hf | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | Confident | A chat-optimized of a 70 billion-parameter model in the Llama 2 series. | Llama 2-70B | llama-2-70b | 113.21 | 109.61 | 116.25 | |||
| Stable Beluga 2 | StableBeluga2 | Stability AI | United Kingdom of Great Britain and Northern Ireland | 2023-07-20 | Open weights (non-commercial) | 2023-07-20 | Likely | Stable Beluga 2 | Stable Beluga 2 | stable-beluga-2 | 116.54 | 112.53 | 119.85 | ||||
| Qwen-1_8B | Qwen-1.8B | 2023-11-30 | Unverified | Qwen's 1.8 billion-parameter model in the original Qwen series. | qwen-1-8b | ||||||||||||
| Qwen-7B | Qwen-7B | Alibaba | China | 2023-09-28 | Open weights (restricted use) | 2023-09-28 | 1.0e+23 | Confident | Qwen's 7 billion-parameter model in the original Qwen series. | Qwen-7B | Qwen-7B | qwen-7b | 105.87 | 98.35 | 110.66 | ||
| Qwen-14B | Qwen-14B | Alibaba | China | 2023-09-28 | Open weights (restricted use) | 2023-09-28 | 2.5e+23 | Confident | Qwen's 14 billion-parameter model in the original Qwen series. | Qwen-14B | Qwen-14B | qwen-14b | 112.30 | 107.92 | 116.96 | ||
| Mistral 7B | Mistral-7B-Instruct-v0.2 | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-10-10 | Confident | Mistral 7B v0.2 | Mistral 7B v0.2 | mistral-7b-v0-2 | |||||||
| Gemma 2B | gemma-2b | Google DeepMind | United States of America | 2024-02-21 | Open weights (restricted use) | 2024-02-21 | 4.5e+22 | Confident | Google's 2 billion-parameter model in the Gemma series of open source models. | Gemma 2B | gemma-2b | 92.81 | 89.05 | 96.29 | |||
| Gemma 7B | gemma-7b | Google DeepMind | United States of America | 2024-02-21 | Open weights (restricted use) | 2024-02-21 | 3.1e+23 | Confident | Google's 7 billion-parameter model in the Gemma series of open source models. | Gemma 7B | gemma-7b | 111.24 | 107.36 | 115.14 | |||
| GPT-3.5 (davinci-003) | text-davinci-003 | 2022-11-28 | 2022-11-28 | Likely | davinci-003 | davinci-003 | davinci-003 | ||||||||||
| GPT-3.5 (davinci-002) | text-davinci-002 | OpenAI | United States of America | 2022-03-15 | API access | 2022-03-15 | 2.6e+24 | Speculative | davinci-002 | davinci-002 | davinci-002 | ||||||
| Mistral 7B | Mistral-7B-v0.1 | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | Confident | Mistral 7B v0.1 | Mistral 7B v0.1 | mistral-7b-v0-1 | 111.46 | 108.26 | 114.19 | ||||
| GPT-3.5 Turbo | gpt-3.5-turbo-0613 | GPT-3.5 Turbo (June 2023) | OpenAI | United States of America | 2023-06-13 | API access | 2023-06-13 | Likely | GPT-3.5 Turbo (June 2023) | GPT-3.5 Turbo (Jun 2023) | GPT-3.5 Turbo (Jun 2023) | gpt-3-5-turbo-jun-2023 | 112.64 | 108.24 | 120.93 | ||
| MPT-30B | mpt-30b-instruct | MosaicML | United States of America | 2023-06-22 | Open weights (unrestricted) | 2023-06-22 | 1.9e+23 | Confident | MPT-30B | mpt-30b | 99.49 | 96.77 | 102.80 | ||||
| Falcon-40B | falcon-40b | Technology Innovation Institute | United Arab Emirates | 2023-03-15 | Open weights (unrestricted) | 2023-03-15 | 2.4e+23 | Confident | The 40 billion-parameter model in the Falcon series. | Falcon-40B | falcon-40b | 103.43 | 100.01 | 106.84 | |||
| Vicuna-13B-v1.3 | vicuna-13b-v1.3 | Large Model Systems Organization,University of California (UC) Berkeley | United States of America | 2023-06-18 | Open weights (restricted use) | 2023-06-22 | Confident | Vicuna-13B-v1.3 | Vicuna-13B-v1.3 | vicuna-13b-v1-3 | |||||||
| OPT-175B | opt-175b | Meta AI | United States of America | 2022-05-02 | Open weights (non-commercial) | 2022-05-02 | 4.3e+23 | Confident | Facebook's 66 billion-parameter model in the OPT series. | OPT-175B | opt-175b | ||||||
| T5-11B | T5-11B | United States of America | Open weights (unrestricted) | 2019-10-23 | 3.3e+22 | Confident | The largest, 11 billion-parameter version of T5, an early large language model from Google. | T5-11B | t5-11b | ||||||||
| OPT-66B | opt-66b | Meta AI | United States of America | 2022-05-03 | Open weights (non-commercial) | 2022-06-21 | 1.1e+23 | Confident | Facebook's 66 billion-parameter model in the OPT series. | OPT-66B | opt-66b | ||||||
| GPT-3 175B (davinci) | davinci | OpenAI | United States of America | API access | 2020-05-28 | 3.1e+23 | Confident | GPT-3 175B (davinci) | GPT-3 175B (davinci) | gpt-3-175b-davinci | |||||||
| BLOOM-176B | bloom | Hugging Face,BigScience | United States of America,France | 2022-07-06 | Open weights (restricted use) | 2022-07-11 | 3.7e+23 | Confident | BLOOM-176B | bloom-176b | |||||||
| MPT-30B | mpt-30b | MosaicML | United States of America | 2023-06-22 | Open weights (unrestricted) | 2023-06-22 | 1.9e+23 | Confident | MPT-30B | mpt-30b | 99.49 | 96.77 | 102.80 | ||||
| GPT-3 6.7B | curie | OpenAI | United States of America | 2020-06-22 | 1.2e+22 | Confident | GPT-3 6.7B | gpt-3-6-7b | |||||||||
| InstructGPT 6B | text-curie-001 | OpenAI | United States of America | API access | 2022-01-27 | Confident | InstructGPT 6B | instructgpt-6b | |||||||||
| GPT-3 Medium | ada | OpenAI | United States of America | 2020-06-22 | 6.4e+20 | Confident | GPT-3 Medium | gpt-3-medium | |||||||||
| GPT-3 XL | babbage | OpenAI | United States of America | 2020-06-22 | 2020-06-22 | 2.4e+21 | Confident | GPT-3 XL (babbage) | GPT-3 XL (babbage) | gpt-3-xl-babbage | |||||||
| InstructGPT 350M | text-ada-001 | OpenAI | United States of America | 2022-01-27 | Confident | InstructGPT 350M | instructgpt-350m | ||||||||||
| InstructGPT 1.3B | text-babbage-001 | OpenAI | United States of America | API access | 2022-01-27 | Confident | InstructGPT 1.3B | instructgpt-1-3b | |||||||||
| InstructGPT 175B | text-davinci-001 | OpenAI | United States of America | 2022-01-27 | API access | 2022-01-27 | 3.2e+23 | Confident | InstructGPT 175B | instructgpt-175b | |||||||
| phi-3.5-mini | Phi-3.5-mini-instruct | Microsoft | United States of America | 2024-08-16 | Open weights (unrestricted) | 2024-04-23 | 3.7e+22 | Confident | An instruction-tuned version of phi-3.5-mini, a small model in Microsoft's Phi 3.5 open source model series. | phi-3.5-mini | phi-3-5-mini | ||||||
| Phi-3.5-MoE | Phi-3.5-MoE-instruct | Microsoft | United States of America | 2024-08-17 | Open weights (unrestricted) | 2024-04-23 | 3.0e+23 | Confident | Phi-3.5-MoE | phi-3-5-moe | |||||||
| Mistral NeMo | Mistral-Nemo-Base-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | Confident | Mistral NeMo | mistral-nemo | 118.13 | 112.83 | 125.53 | |||||
| Gemma 2 9B | gemma-2-9b | Google DeepMind | United States of America | Open weights (restricted use) | 2024-06-24 | 4.3e+23 | Confident | A 9 billion-parameter open source model from Google DeepMind in the Gemma 2 series. | Gemma 2 9B | gemma-2-9b | 119.35 | 115.40 | 122.93 | ||||
| Megatron-Turing NLG 530B | Megatron-Turing NLG 530B | Microsoft,NVIDIA | United States of America | 2022-01-28 | Unreleased | 2021-10-11 | 8.6e+23 | Confident | Megatron-Turing NLG 530B | megatron-turing-nlg-530b | |||||||
| Llama 2-34B | Llama-2-34b | Meta AI | United States of America | 2023-07-18 | Unreleased | 2023-07-18 | 4.1e+23 | Confident | Meta's 34 billion-parameter model in the Llama 2 series. | Llama 2-34B | llama-2-34b | 104.21 | 100.63 | 109.34 | |||
| Gopher (280B) | Gopher (280B) | DeepMind | United Kingdom of Great Britain and Northern Ireland | 2021-12-08 | Unreleased | 2021-12-08 | 6.3e+23 | Confident | A 280 billion-parameter model from DeepMind in 2021. | Gopher (280B) | gopher-280b | ||||||
| Chinchilla | Chinchilla (70B) | DeepMind | United Kingdom of Great Britain and Northern Ireland | 2022-03-29 | Unreleased | 2022-03-29 | 5.8e+23 | Confident | Chinchilla | Chinchilla | chinchilla | ||||||
| PaLM 62B | 2022-04-04 | Unverified | Google's 62 billion-parameter model in the PaLM series. | palm-62b | |||||||||||||
| PaLM (540B) | PaLM 540B | Google Research | United States of America | 2022-04-04 | Unreleased | 2022-04-04 | 2.5e+24 | Confident | Google's largest model in the PaLM series. | PaLM (540B) | palm-540b | ||||||
| vicuna-13b-v1.1 | 2023-04-12 | Unverified | vicuna-13b-v1-1 | ||||||||||||||
| OPT-1.3B | opt-1.3b | Meta AI | United States of America | 2022-05-11 | Open weights (non-commercial) | 2022-06-21 | Confident | OPT-1.3B | opt-1-3b | ||||||||
| GPT-Neo-2.7B | gpt-neo-2.7B | EleutherAI | United States of America | 2023-03-30 | Open weights (unrestricted) | 2021-03-21 | 7.9e+21 | Confident | GPT-Neo-2.7B | gpt-neo-2-7b | |||||||
| GPT-2 (1.5B) | gpt2-xl | OpenAI | United States of America | 2019-11-05 | Open weights (unrestricted) | 2019-02-14 | 1.9e+21 | Speculative | The largest version of GPT-2, OpenAI's second generation transformer-based large pretrained language model. | GPT-2 (1.5B) | gpt-2-1-5b | ||||||
| Phi-1.5 | phi-1_5 | Microsoft | United States of America | 2023-09-11 | Open weights (unrestricted) | 2023-09-11 | 1.2e+21 | Confident | Phi-1.5 | phi-1-5 | 90.07 | 67.64 | 97.77 | ||||
| PaLM 2-S | 2023-05-17 | Unverified | DeepMind's smallest model in the PaLM 2 series. | palm-2-s | |||||||||||||
| PaLM 2-M | 2023-05-17 | Unverified | DeepMind's mid-sized model in the PaLM 2 series. | palm-2-m | |||||||||||||
| PaLM 2-L | 2023-05-17 | Unverified | DeepMind's largest model in the PaLM 2 series. | palm-2-l | |||||||||||||
| Falcon-180B | falcon-180B | Technology Innovation Institute | United Arab Emirates | 2023-09-06 | Open weights (restricted use) | 2023-09-06 | 3.8e+24 | Confident | The 180 billion-parameter model in the Falcon series. | Falcon-180B | falcon-180b | 111.40 | 107.68 | 123.77 | |||
| Inflection-1 | Inflection-1 | Inflection AI | United States of America | 2023-06-22 | Hosted access (no API) | 2023-06-23 | 1.0e+24 | Speculative | Inflection-1 | inflection-1 | |||||||
| XGen-7B | xgen-7b-8k-base | Salesforce | United States of America | 2023-06-27 | Open weights (unrestricted) | 2023-09-07 | 8.0e+22 | Confident | XGen-7B | XGen-7B | xgen-7b | 91.99 | 88.28 | 94.79 | |||
| open_llama_7b | 2023-06-07 | Unverified | open-llama-7b | ||||||||||||||
| RedPajama-INCITE-7B-Base | 2023-05-04 | Unverified | redpajama-incite-7b-base | ||||||||||||||
| GPT-NeoX-20B | gpt-neox-20b | EleutherAI | United States of America | 2022-04-07 | Open weights (unrestricted) | 2022-02-09 | 9.3e+22 | Confident | GPT-NeoX-20B | gpt-neox-20b | |||||||
| opt-13b | 2022-05-11 | Unverified | opt-13b | ||||||||||||||
| GPT-J-6B | gpt-j-6b | EleutherAI,LAION | United States of America,Germany | 2021-08-05 | Open weights (unrestricted) | 2021-05-01 | 1.5e+22 | Confident | GPT-J-6B | gpt-j-6b | |||||||
| Dolly 2.0-12b | dolly-v2-12b | Databricks | United States of America | 2023-04-11 | Open weights (unrestricted) | 2023-04-12 | Confident | Dolly 2.0-12b | dolly-2-0-12b | 88.12 | 80.39 | 93.66 | |||||
| Cerebras-GPT-13B | Cerebras-GPT-13B | Cerebras Systems | United States of America | 2023-03-20 | Open weights (unrestricted) | 2023-04-06 | 2.3e+22 | Confident | Cerebras-GPT-13B | cerebras-gpt-13b | 81.49 | 72.99 | 85.20 | ||||
| stablelm-tuned-alpha-7b | 2023-04-19 | Unverified | stablelm-tuned-alpha-7b | ||||||||||||||
| T5-Small | Unverified | The smallest version of T5, an early large language model from Google. | t5-small | ||||||||||||||
| T5-Base | Unverified | A version of T5, an early large language model from Google. | t5-base | ||||||||||||||
| T5-Large | Unverified | A version of T5, an early large language model from Google. | t5-large | ||||||||||||||
| T5-3B | T5-3B | United States of America | Open weights (unrestricted) | 2019-10-23 | 9.0e+21 | Confident | The 3 billion-parameter version of T5, an early large language model from Google. | T5-3B | t5-3b | ||||||||
| GPT-5.4 Pro | gpt-5.4-pro-2026-03-05_unknown | GPT-5.4 Pro | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 Pro | GPT-5.4 Pro | gpt-5-4-pro | 157.91 | 155.37 | 165.49 | |||
| GPT-5 | gpt-5-2025-08-07_unknown | GPT-5 (unknown thinking) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 (high) | GPT-5 | gpt-5 | 150.00 | 147.33 | 156.34 | |
| GPT-5 mini | gpt-5-mini-2025-08-07_unknown | GPT-5 mini (unknown thinking) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 mini (high) | GPT-5 mini | gpt-5-mini | 145.83 | 143.09 | 151.48 | ||
| o1-pro | o1-pro-2025-03-19 | o1 Pro | OpenAI | United States of America | 2025-03-19 | API access | 2025-03-19 | Likely | An An enhanced version of OpenAI's first reasoning model,o1, evaluated with low reasoning effort. | o1-pro | o1-pro | ||||||
| o1 | o1-2024-12-17_unknown | o1 (unknown thinking) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | Unknown | OpenAI's first reasoning model, evaluated with at the high reasoning effort level. | o1 | o1 | 142.74 | 140.13 | 147.88 | |||
| xiaoyi-deepresearch | 2026-03-06 | Unverified | xiaoyi-deepresearch | ||||||||||||||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_2K | Claude Sonnet 4.5 (2k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (2k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 146.85 | 144.23 | 152.58 | |||
| Claude Opus 4 | claude-opus-4-20250514_2K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 2,000 reasoning tokens. | Claude Opus 4 (2k thinking) | Claude Opus 4 | claude-opus-4 | 143.10 | 140.42 | 148.78 | ||
| Claude Sonnet 4 | claude-sonnet-4-20250514_2K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized Claude 4 series model, evaluated at up to 2,000 reasoning tokens. | Claude Sonnet 4 (2k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.32 | 139.26 | 147.36 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_2K | Claude 3.7 Sonnet (2k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 2,000 reasoning tokens allowed. | Claude 3.7 Sonnet (2k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 141.83 | 139.34 | 147.11 | |
| Sonar | sonar-pro | Perplexity | United States of America | 2025-01-21 | API access | 2025-02-11 | Confident | A search-optimized language model from Perplexity. | Sonar Pro | Sonar | sonar | ||||||
| Baichuan1-7B | Baichuan-7B | Baichuan | China | 2023-06-01 | Open weights (non-commercial) | 2023-06-01 | 5.0e+22 | Confident | Baichuan1-7B | baichuan1-7b | 88.92 | 82.72 | 95.56 | ||||
| DeepSeek-V2 (MoE-236B) | DeepSeek-V2 | DeepSeek | China | 2024-05-07 | Open weights (restricted use) | 2024-05-07 | 1.0e+24 | Confident | DeepSeek-V2 (MoE-236B, May 2024) | DeepSeek-V2 (MoE-236B, May 2024) | deepseek-v2-moe-236b-may-2024 | 124.46 | 121.32 | 128.75 | |||
| Qwen2.5-72B | Qwen2.5-72B | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | Confident | A 72 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-72B | Qwen2.5-72B | qwen2-5-72b | 129.08 | 123.83 | 133.37 | ||
| Llama 3.1-405B | Llama-3.1-405B | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 3.8e+25 | Confident | The 405 billion-parameter model in Meta’s LLaMA 3.1 series. | Llama 3.1-405B | llama-3-1-405b | 128.99 | 126.57 | 132.94 | |||
| DeepSeek-V3 | DeepSeek-V3-Base | DeepSeek | China | 2024-12-26 | Open weights (restricted use) | 2024-12-24 | 3.3e+24 | Confident | DeepSeek-V3 (base) | DeepSeek-V3 | deepseek-v3 | 133.03 | 129.36 | 139.00 | |||
| Mistral 7B | Mistral-7B-Instruct-v0.1 | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | Confident | Mistral 7B v0.1 | Mistral 7B v0.1 | mistral-7b-v0-1 | 111.46 | 108.26 | 114.19 | ||||
| Mixtral 8x7B | Mixtral-8x7B-v0.1 | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | Speculative | Mixtral 8x7B | mixtral-8x7b | 117.84 | 114.43 | 120.59 | ||||
| Nemotron-4 15B | Nemotron-4 15B | NVIDIA | United States of America | 2024-02-26 | Unreleased | 2024-02-27 | 7.5e+23 | Confident | A 15 billion-parameter language model trained by Nvidia. | Nemotron-4 15B | nemotron-4-15b | 106.81 | 103.48 | 111.75 | |||
| Falcon 2 11B | falcon-11b | Falcon 2-11B | Technology Innovation Institute | United Arab Emirates | 2024-05-09 | Open weights (restricted use) | 2024-05-09 | 3.6e+23 | Confident | Falcon 2 11B | falcon-2-11b | 108.70 | 105.64 | 115.66 | |||
| GLaM | GLaM (MoE) | United States of America | 2021-12-13 | Unreleased | 2021-12-13 | 3.6e+23 | Confident | GLaM | glam | ||||||||
| GPT-4 (Mar 2023) | gpt-4-32k-0314 | OpenAI | United States of America | 2023-03-14 | API access | 2023-03-15 | 2.1e+25 | Likely | GPT-4 (Mar 2023) | gpt-4-mar-2023 | 125.82 | 121.38 | 133.35 | ||||
| INTELLECT-1 | INTELLECT-1-Instruct | Prime Intellect,Hugging Face,Arcee AI | United States of America | 2024-11-29 | 2024-11-29 | 6.0e+22 | Confident | INTELLECT-1 | intellect-1 | 99.70 | 95.81 | 102.44 | |||||
| Phi-2 | phi-2 | Microsoft | United States of America | 2023-12-12 | Open weights (unrestricted) | 2023-12-12 | 2.3e+22 | Confident | Phi-2 | phi-2 | 107.01 | 94.28 | 111.78 | ||||
| Qwen2.5-Coder-0.5B | 2024-09-18 | Unverified | Qwen's 500 million-parameter model in the Qwen 2.5 Coder series of open source models. | qwen2-5-coder-0-5b | |||||||||||||
| Qwen2.5-Coder (1.5B) | Qwen2.5-Coder-1.5B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-09-18 | 5.1e+22 | Confident | Qwen's 1.5 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-1.5B | Qwen2.5-Coder (1.5B) | qwen2-5-coder-1-5b | 101.90 | 86.61 | 108.33 | ||
| Qwen2.5-Coder-3B | 2024-09-18 | Unverified | Qwen's 3 billion-parameter model in the Qwen 2.5 Coder series of open source models. | qwen2-5-coder-3b | |||||||||||||
| Qwen2.5-Coder (7B) | Qwen2.5-Coder-7B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-09-18 | 2.5e+23 | Confident | Qwen's 7 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-7B | Qwen2.5-Coder (7B) | qwen2-5-coder-7b | 112.42 | 104.98 | 118.79 | ||
| Qwen2.5-Coder-14B | 2024-09-18 | Unverified | Qwen's 14 billion-parameter model in the Qwen 2.5 Coder series of open source models. | qwen2-5-coder-14b | |||||||||||||
| Qwen2.5-Coder (32B) | Qwen2.5-Coder-32B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-11-12 | 1.1e+24 | Confident | Qwen's 32 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-32B | Qwen2.5-Coder-32B | Qwen2.5-Coder-32B | qwen2-5-coder-32b | 118.98 | 113.03 | 125.64 | |
| Yi 6B | Yi-6B | 01.AI | China | 2023-11-02 | Open weights (restricted use) | 2023-11-02 | 1.3e+23 | Confident | Yi-6B (chat) | Yi 6B | yi-6b | 103.75 | 100.09 | 109.21 | |||
| Yi-9B | 2024-03-01 | Unverified | yi-9b | ||||||||||||||
| GPT-5.5 Pro | gpt-5.5-pro_unknown | GPT-5.5 Pro (unknown thinking) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 Pro | gpt-5-5-pro | 161.29 | 158.43 | 170.30 | ||||
| Claude Opus 4 | claude-opus-4-20250514_12K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 12,000 reasoning tokens. | Claude Opus 4 (12k thinking) | Claude Opus 4 | claude-opus-4 | 143.10 | 140.42 | 148.78 | ||
| GPT-5.2 Pro | gpt-5.2-pro-2025-12-11 | GPT-5.2 Pro | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Likely | GPT-5.2 Pro | GPT-5.2 Pro | gpt-5-2-pro | 154.72 | 152.20 | 161.41 | |||
| Grok 4.1 Fast | grok-4-1-fast-non-reasoning | xAI | United States of America | 2025-11-19 | API access | 2025-11-19 | Unknown | Grok 4.1 Fast | grok-4-1-fast | ||||||||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_12K | Claude Sonnet 4.5 (12k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (12k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 146.85 | 144.23 | 152.58 | |||
| DeepSeek-V3.2 | DeepSeek-V3.2-Speciale | DeepSeek | China | 2025-12-01 | Open weights (unrestricted) | 2025-12-01 | 4.2e+24 | Likely | DeepSeek-V3.2-Speciale | DeepSeek-V3.2-Speciale | deepseek-v3-2-speciale | ||||||
| GLM-4.7 | zai-org/glm-4.7 | GLM-4.7 (Novita) | Z.ai (Zhipu AI) | China | 2025-12-22 | Open weights (unrestricted) | 2025-12-22 | 4.4e+24 | Likely | GLM-4.7 | glm-4-7 | 143.84 | 139.56 | 149.42 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_12K | Claude 3.7 Sonnet (12k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 12,000 reasoning tokens allowed. | Claude 3.7 Sonnet (12k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 141.83 | 139.34 | 147.11 | |
| Claude Sonnet 4 | claude-sonnet-4-20250514_12K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized Claude 4 series model, evaluated at up to 12,000 reasoning tokens. | Claude Sonnet 4 (12k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.32 | 139.26 | 147.36 | |||
| DeepSeek-V3.1 | DeepSeek-V3.1 | DeepSeek | China | 2025-08-21 | Open weights (unrestricted) | 2025-08-21 | 3.6e+24 | Confident | An updated version of DeepSeek V3. | DeepSeek-V3.1 | DeepSeek-V3.1 | deepseek-v3-1 | 138.75 | 135.48 | 145.60 | ||
| Baichuan 1-13B | Baichuan-13B-Base | Baichuan | China | 2023-07-11 | Open weights (restricted use) | 2023-07-11 | 9.4e+22 | Confident | Baichuan 1-13B | baichuan-1-13b | |||||||
| Baichuan2-13B-Chat | 2023-09-06 | Unverified | A chat-optimized version of Baichuan's 13 billion-parameter model in the Baichuan 2 series. | baichuan2-13b-chat | |||||||||||||
| Claude 1.3 | claude-1.3 | Anthropic | United States of America | 2023-04-18 | API access | 2023-04-18 | Unknown | An updated version of Anthropic's first generation Claude model. | Claude 1.3 | claude-1-3 | |||||||
| Claude Instant | claude-instant-1.1 | Anthropic | United States of America | API access | 2023-08-09 | Unknown | Claude Instant | claude-instant | 120.88 | 117.48 | 125.04 | ||||||
| Claude Instant | claude-instant-1.2 | Anthropic | United States of America | 2023-08-09 | API access | 2023-08-09 | Unknown | Claude Instant | claude-instant | 120.88 | 117.48 | 125.04 | |||||
| CodeQwen1.5-7B | 2024-04-15 | Unverified | codeqwen1-5-7b | ||||||||||||||
| DeepSeek Coder 1.3B | deepseek-coder-1.3b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 1.6e+22 | Likely | DeepSeek Coder 1.3B (base) | DeepSeek Coder 1.3B | deepseek-coder-1-3b | 61.93 | 57.26 | 77.78 | |||
| DeepSeek Coder 6.7B | deepseek-coder-6.7b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 8.0e+22 | Likely | DeepSeek Coder 6.7B (base) | DeepSeek Coder 6.7B | deepseek-coder-6-7b | 88.07 | 80.46 | 97.38 | |||
| DeepSeek Coder 33B | deepseek-coder-33b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 4.0e+23 | Likely | DeepSeek Coder 33B (base) | DeepSeek Coder 33B | deepseek-coder-33b | 94.99 | 88.78 | 101.53 | |||
| DeepSeek-Coder-V2-Lite-Base | 2024-06-13 | Unverified | deepseek-coder-v2-lite-base | ||||||||||||||
| Gemini 1.5 Flash | gemini-1.5-flash-0514 | Google DeepMind | United States of America | 2024-05-14 | API access | 2024-05-10 | Unknown | Gemini 1.5 Flash (May 2024) | Gemini 1.5 Flash (May 2024) | gemini-1-5-flash-may-2024 | 122.32 | 118.95 | 125.93 | ||||
| GPT-4 Turbo (Nov 2023) | gpt-4-turbo | OpenAI | United States of America | 2023-11-06 | API access | 2023-11-06 | Unknown | An updated version of GPT-4 and OpenAI's then-new flagship language model. | GPT-4 Turbo (Nov 2023) | gpt-4-turbo-nov-2023 | |||||||
| internlm-chat-20b | 2023-09-17 | Unverified | internlm-chat-20b | ||||||||||||||
| Llama 3.2 11B | Llama-3.2-11B-Vision-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 5.8e+23 | Confident | An 11 billion-parameter instruction-tuned vision language model in Meta's Llama 3.2 series. | Llama 3.2 11B | llama-3-2-11b | ||||||
| Llama 2-13B | Llama-2-13b-chat | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 1.6e+23 | Confident | A chat-optimized version of Meta's 13 billion-parameter model in the Llama 2 series. | Llama 2-13B | llama-2-13b | 105.20 | 102.56 | 108.98 | |||
| Llama 2-70B | Llama-2-70b-chat | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | Confident | A chat-optimized version of Meta's 70 billion-parameter model in the Llama 2 series. | Llama 2-70B | llama-2-70b | 113.21 | 109.61 | 116.25 | |||
| mistral-small-2402 | 2024-02-26 | Unverified | A 2024 version of a Mistral language model. | mistral-small-2402 | |||||||||||||
| Mixtral 8x22B | Mixtral-8x22B-v0.1 | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | Speculative | Mixtral 8x22B | mixtral-8x22b | 121.08 | 114.24 | 124.41 | ||||
| Qwen-14B | Qwen-14B-Chat | Alibaba | China | 2023-09-24 | Open weights (restricted use) | 2023-09-28 | 2.5e+23 | Confident | A chat-optimized version of Qwen's first-generation 14 billion-parameter model. | Qwen-14B (chat) | Qwen-14B | qwen-14b | 112.30 | 107.92 | 116.96 | ||
| Qwen1.5-7B | Qwen1.5-7B | Alibaba | China | 2024-02-04 | Open weights (unrestricted) | 2024-02-04 | 1.7e+23 | Confident | A 7 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-7B | Qwen1.5-7B | qwen1-5-7b | |||||
| Qwen1.5-14B | qwen1.5-14B | Alibaba | China | 2024-02-04 | Open weights (unrestricted) | 2024-02-04 | 3.4e+23 | Confident | A 14 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-14B | Qwen1.5-14B | qwen1-5-14b | |||||
| Qwen1.5-32B | qwen1.5-32B | Alibaba | China | 2024-02-04 | Open weights (restricted use) | 2024-02-05 | Confident | A 32 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-32B | Qwen1.5-32B | qwen1-5-32b | ||||||
| Qwen2.5-14B | qwen2.5-14b-instruct | Alibaba | China | 2025-02-26 | Open weights (unrestricted) | 2024-09-19 | 1.6e+24 | Confident | An instruction-tuned version of Qwen 2.5 14b. | Qwen2.5-14B (instruct) | Qwen2.5-14B | qwen2-5-14b | |||||
| Qwen2.5-7B | qwen2.5-7b-instruct | Alibaba | China | 2025-02-26 | Open weights (unrestricted) | 2024-09-19 | 8.2e+23 | Confident | An instruction-tuned version of Qwen 2.5 7b. | Qwen2.5-7B | Qwen2.5-7B | qwen2-5-7b | |||||
| StarCoder 2 3B | starcoder2-3b | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | 2024-02-22 | Open weights (restricted use) | 2024-02-29 | 5.9e+22 | Confident | StarCoder 2 3B | StarCoder 2 3B | starcoder-2-3b | 87.19 | 78.18 | 91.96 | |||
| StarCoder 2 7B | starcoder2-7b | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | 2024-02-20 | Open weights (restricted use) | 2024-02-29 | 1.5e+23 | Confident | StarCoder 2 7B | StarCoder 2 7B | starcoder-2-7b | 92.13 | 79.14 | 97.88 | |||
| StarCoder 2 15B | starcoder2-15b | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | 2024-02-20 | Open weights (restricted use) | 2024-02-29 | 3.9e+23 | Confident | StarCoder 2 15B | StarCoder 2 15B | starcoder-2-15b | 103.99 | 97.38 | 110.87 | |||
| Yi-34B | Yi-34B | 01.AI | China | 2023-11-02 | Open weights (restricted use) | 2023-11-02 | 6.1e+23 | Confident | Yi-34B | Yi-34B | Yi-34B | yi-34b | 116.82 | 111.91 | 120.05 | ||
| Yi-Large | Yi-large | 01.AI | China | 2024-05-13 | API access | 2024-05-13 | 1.8e+24 | Speculative | Yi-Large | Yi-Large | yi-large | ||||||
| Yi 6B | Yi-6B-Chat | 01.AI | China | 2023-11-22 | Open weights (restricted use) | 2023-11-02 | 1.3e+23 | Confident | Yi-6B | Yi 6B | yi-6b | 103.75 | 100.09 | 109.21 | |||
| o3 | o3-2025-04-16_unknown | o3 (unknown thinking) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | Unknown | OpenAI's second generation o-series reasoning model, evaluated with high reasoning effort. | o3 | o3 | 147.10 | 144.35 | 152.96 | |||
| o4-mini | o4-mini-2025-04-16_unknown | o4-mini (unknown thinking) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | Unknown | The latest smaller size o-series reasoning model from OpenAI, evaluated with high reasoning effort. | o4-mini | o4-mini | 146.53 | 143.44 | 153.20 | |||
| Kimi K2 | Kimi-K2-Instruct-0905 | Moonshot | China | 2025-09-05 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | Kimi K2 (Sep 2025) | Kimi K2 (Sep 2025) | kimi-k2-sep-2025 | 141.04 | 137.75 | 147.36 | |||
| Qwen3-235B-A22B | Qwen3-235B-A22B | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Likely | Qwen3 (Apr 2025) | Qwen3-235B-A22B | qwen3-235b-a22b | 139.61 | 134.81 | 144.40 | |||
| o3-mini | o3-mini-2025-01-31_unknown | o3-mini (unknown thinking) | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | Unknown | A smaller version of OpenAI's o3 reasoning model, evaluated with high reasoning effort. | o3-mini | o3-mini | 141.40 | 138.74 | 146.89 | |||
| Claude Sonnet 4.6 | claude-sonnet-4-6_unknown | Claude Sonnet 4.6 (unknown thinking) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 153.41 | 150.31 | 160.16 | ||||
| GPT-5 nano | gpt-5-nano-2025-08-07_unknown | GPT-5 nano (unknown thinking) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 nano (high) | GPT-5 nano | gpt-5-nano | 140.85 | 136.04 | 146.03 | ||
| Qwen1.5-110B | qwen1.5-110b-chat | Alibaba | China | 2024-04-25 | Open weights (unrestricted) | 2024-04-25 | Confident | Qwen1.5-110B | Qwen1.5-110B | qwen1-5-110b | |||||||
| GPT-5.4 Nano | gpt-5.4-nano-2026-03-17_unknown | GPT-5.4 nano (unknown thinking) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 nano (unknown thinking) | GPT-5.4 Nano | gpt-5-4-nano | 146.54 | 143.59 | 152.29 | |||
| Llama 3-70B | Meta-Llama-3-70B | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.9e+24 | Confident | Meta's 70 billion-parameter model in the Llama 3 series. | Llama 3-70B | llama-3-70b | 122.47 | 120.41 | 126.46 | |||
| Gemini 2.0 Flash-Lite | gemini-2.0-flash-lite-001 | Google DeepMind | United States of America | 2025-02-25 | API access | 2024-02-05 | Unknown | The first version of the smallest model in Google DeepMind's Gemini 2.0 series. | Gemini 2.0 Flash-Lite | gemini-2-0-flash-lite | |||||||
| GPT-5.4 Mini | gpt-5.4-mini-2026-03-17_unknown | GPT-5.4 mini (unknown thinking) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 mini (unknown thinking) | GPT-5.4 Mini | gpt-5-4-mini | 148.96 | 145.80 | 155.00 | |||
| Mixtral 8x22B | Mixtral-8x22B-Instruct-v0.1 | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | Speculative | An instruction-tuned version of Mistral's 8 billion active (176 billion total)-parameter mixture-of-experts model | Mixtral 8x22B | mixtral-8x22b | 121.08 | 114.24 | 124.41 | |||
| Llama 3-8B | Meta-Llama-3-8B | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.2e+23 | Confident | Meta's 8 billion-parameter model in the Llama 3 series. | Llama 3-8B | llama-3-8b | 116.14 | 112.69 | 118.47 | |||
| GPT-3.5 (davinci-002) | code-davinci-002 | OpenAI | United States of America | 2022-03-15 | API access | 2022-03-15 | 2.6e+24 | Speculative | davinci-002 | davinci-002 | davinci-002 | ||||||
| Falcon-40B | falcon-40b-instruct | Technology Innovation Institute | United Arab Emirates | 2023-05-25 | Open weights (unrestricted) | 2023-03-15 | 2.4e+23 | Confident | Falcon-40B | falcon-40b | 103.43 | 100.01 | 106.84 | ||||
| DeepSeek-Coder-V2-Lite-Instruct | 2024-06-13 | Unverified | deepseek-coder-v2-lite-instruct | ||||||||||||||
| DeepSeek-Coder-V2 236B | DeepSeek-Coder-V2-Instruct | DeepSeek | China | 2024-06-17 | Open weights (restricted use) | 2024-06-17 | 1.3e+24 | Confident | An instruction-tuned version of DeepSeek Coder V2, a 236 billion-parameter coding-optimized model from DeepSeek. | DeepSeek-Coder-V2 (instruct) | DeepSeek-Coder-V2 236B | deepseek-coder-v2-236b | |||||
| Qwen2.5-Coder-3B-Instruct | 2024-11-06 | Unverified | An instruction-tuned and coding-optimized 3 billion-parameter model in the Qwen 2.5 series. | qwen2-5-coder-3b-instruct | |||||||||||||
| Qwen2.5-Coder (7B) | Qwen2.5-Coder-7B-Instruct | Alibaba | China | 2024-09-17 | Open weights (unrestricted) | 2024-09-18 | 2.5e+23 | Confident | An instruction-tuned and coding-optimized 7 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-Coder-7B (instruct) | Qwen2.5-Coder (7B) | qwen2-5-coder-7b | 112.42 | 104.98 | 118.79 | ||
| Qwen2.5-Coder-14B-Instruct | 2024-11-06 | Unverified | An instruction-tuned and coding-optimized 14 billion-parameter model in the Qwen 2.5 series. | qwen2-5-coder-14b-instruct | |||||||||||||
| DeepSeek-Coder-V2 236B | DeepSeek-Coder-V2-Base | DeepSeek | China | 2024-06-17 | Open weights (restricted use) | 2024-06-17 | 1.3e+24 | Confident | DeepSeek-Coder-V2 (base) | DeepSeek-Coder-V2 236B | deepseek-coder-v2-236b | ||||||
| BLIP-2 (Q-Former) | blip2-opt-2.7b | Salesforce Research | United States of America | 2023-02-06 | Open weights (unrestricted) | 2023-01-30 | 1.2e+21 | Confident | BLIP-2 (Q-Former) | blip-2-q-former | |||||||
| Phi-3.5-vision-instruct | 2024-08-16 | Unverified | An multimodal version in Microsoft's Phi-3.5 series of open source models. | phi-3-5-vision-instruct | |||||||||||||
| MM1-3B-Chat | 2024-03-14 | Unverified | mm1-3b-chat | ||||||||||||||
| MM1-7B-Chat | 2024-03-14 | Unverified | mm1-7b-chat | ||||||||||||||
| llava-v1.6-vicuna-7b | 2024-01-31 | Unverified | llava-v1-6-vicuna-7b | ||||||||||||||
| llama3-llava-next-8b | 2024-04-20 | Unverified | llama3-llava-next-8b | ||||||||||||||
| Gemini 1.0 Pro Vision | gemini-1.0-pro-vision | Google DeepMind | United States of America | 2024-01-04 | API access | 2024-02-15 | Unknown | A vision language model from Google DeepMind. | Gemini 1.0 Pro Vision | gemini-1-0-pro-vision | |||||||
| falcon-11B-vlm | 2024-05-21 | Unverified | falcon-11b-vlm | ||||||||||||||
| llava-v1.6-vicuna-13b | 2024-01-31 | Unverified | llava-v1-6-vicuna-13b | ||||||||||||||
| llava-v1.6-mistral-7b | 2024-01-31 | Unverified | llava-v1-6-mistral-7b | ||||||||||||||
| instructblip-vicuna-7b | 2023-05-22 | Unverified | instructblip-vicuna-7b | ||||||||||||||
| instructblip-vicuna-13b | 2023-12-25 | Unverified | instructblip-vicuna-13b | ||||||||||||||
| InternVL-Chat-ViT-6B-Vicuna-7B | 2023-12-25 | Unverified | internvl-chat-vit-6b-vicuna-7b | ||||||||||||||
| InternVL-Chat-ViT-6B-Vicuna-13B | 2024-08-16 | Unverified | internvl-chat-vit-6b-vicuna-13b | ||||||||||||||
| llava-v1.5-7b | 2023-10-05 | Unverified | llava-v1-5-7b | ||||||||||||||
| DeepSeek-V4-Flash | deepseek-v4-flash_max | DeepSeek v4 Flash (max) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 2.5e+24 | Likely | DeepSeek v4 (max) | DeepSeek-V4-Flash | deepseek-v4-flash | |||||
| Nemotron 3 Ultra | nemotron-3-ultra | NVIDIA | United States of America | 2026-06-04 | Open weights (unrestricted) | 2026-06-04 | 6.6e+24 | Likely | Nemotron 3 Ultra | nemotron-3-ultra | |||||||
| gpt-oss-20b | gpt-oss-20b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | Confident | The smaller one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-20b | gpt-oss-20b | ||||||
| Switch-Base | Unverified | switch-base | |||||||||||||||
| Switch-Large | Unverified | switch-large | |||||||||||||||
| Llama 4 Maverick | chutes/Llama-4-Maverick-17B-128E-Instruct | Meta AI | United States of America | 2025-04-06 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | Likely | Llama 4 Maverick (128 experts) (Chutes) | Llama 4 Maverick | llama-4-maverick | 133.04 | 130.31 | 137.25 | |||
| Claude Fable 5 | claude-fable-5 | Claude Fable 5 | Anthropic | United States of America | 2026-06-09 | API access | 2026-06-09 | Confident | Claude Fable 5 | Claude Fable 5 | claude-fable-5 | 161.00 | 158.16 | 169.22 | |||
| GPT‑5-Codex | gpt-5-codex | GPT-5-codex | OpenAI | United States of America | 2025-09-15 | API access | 2025-09-15 | Unknown | GPT-5-codex | GPT‑5-Codex | gpt5-codex | ||||||
| gpt-oss-120b | gpt-oss-120b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | The larger one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-120b | gpt-oss-120b | 140.55 | 135.47 | 146.61 | |||
| Reka Flash 3 | reka-flash-3 | Reka AI | United States of America | 2025-03-10 | Open weights (unrestricted) | 2025-03-10 | Confident | Reka Flash 3 | Reka Flash 3 | reka-flash-3 | |||||||
| Mistral NeMo | Mistral-Nemo-Instruct-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | Confident | Mistral NeMo | mistral-nemo | 118.13 | 112.83 | 125.53 | |||||
| Llama 3.2 3B | Llama-3.2-3B-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 1.7e+23 | Confident | A 3 billion-parameter instruction-tuned model in Meta's Llama 3.2 series. | Llama 3.2 3B | llama-3-2-3b | ||||||
| Llama 3.2 1B | Llama-3.2-1B-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 6.6e+22 | Confident | A 1 billion-parameter instruction-tuned model in Meta's Llama 3.2 series. | Llama 3.2 1B | llama-3-2-1b | ||||||
| openhands-lm-32b-v0.1 | 2024-03-26 | Unverified | openhands-lm-32b-v0-1 | ||||||||||||||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro-preview-06-05_32K | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America | 2025-06-05 | API access | 2025-06-05 | Unknown | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model, evaluated with up to 32,000 reasoning tokens. | Gemini 2.5 Pro Preview (Jun 2025, 32k thinking) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.17 | 143.36 | 151.99 | |
| Claude Opus 4 | claude-opus-4-20250514_32K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 32,000 reasoning tokens. | Claude Opus 4 (32k thinking) | Claude Opus 4 | claude-opus-4 | 143.10 | 140.42 | 148.78 | ||
| Qwen3-32B | Qwen3-32B | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 7.1e+24 | Confident | The 32 billion-parameter model in the Qwen 3 series. | Qwen3 32B | Qwen3-32B | qwen3-32b | |||||
| GPT-4o | chatgpt-4o-03-27-2025 | OpenAI | United States of America | 2025-03-27 | API access | 2024-05-13 | Speculative | A version of GPT-4o that was behind the ChatGPT interface, released in March 2025. | ChatGPT-4o (Mar 2025) | GPT-4o (Mar 2025) | GPT-4o (Mar 2025) | gpt-4o-mar-2025 | |||||
| Cohere Command A | c4ai-command-a-03-2025 | Cohere | Canada | 2025-03-13 | Open weights (non-commercial) | 2025-03-13 | Confident | Command A | Cohere Command A | cohere-command-a | |||||||
| DeepSeek-V2.5 | DeepSeek-V2.5 | DeepSeek | China | 2024-09-06 | Open weights (restricted use) | 2024-09-06 | 1.8e+24 | Confident | DeepSeek-V2.5 (Sept 2024) | DeepSeek-V2.5 (Sep 2024) | DeepSeek-V2.5 (Sep 2024) | deepseek-v2-5-sep-2024 | |||||
| GPT-4o | chatgpt-4o-01-29-2025 | OpenAI | United States of America | 2025-01-29 | API access | 2024-05-13 | Speculative | A version of GPT-4o that was behind the ChatGPT interface, released in January 2025. | ChatGPT-4o (Jan 2025) | GPT-4o (Jan 2025) | GPT-4o (Jan 2025) | gpt-4o-jan-2025 | |||||
| Codestral | codestral-2501 | Mistral AI | France | 2025-01-13 | Open weights (non-commercial) | 2024-05-29 | Confident | A January 2025 version of a 2024 coding-optimized model from Mistral. | Codestral | codestral | |||||||
| Yi-Lightning | yi-lightning | 01.AI | China | 2024-12-02 | API access | 2024-10-18 | 1.5e+24 | Confident | Yi-Lightning | Yi-Lightning | yi-lightning | ||||||
| DeepSeek-V3.2 | deepseek-chat | DeepSeek | China | 2025-12-01 | Open weights (unrestricted) | 2025-12-01 | 4.2e+24 | Likely | DeepSeek-V3.2 | deepseek-v3-2 | 146.64 | 143.60 | 152.74 | ||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-04-17 (24K thinking) | Google DeepMind | United States of America | 2025-04-17 | API access | 2025-04-17 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 24,000 reasoning tokens. | Gemini 2.5 Flash (Apr 2025) | Gemini 2.5 Flash (Apr 2025) | gemini-2-5-flash-apr-2025 | 140.55 | 136.89 | 146.50 | |||
| o1 | o1-2024-12-17_low | o1 (low) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | Unknown | OpenAI's first reasoning model, evaluated with at the low reasoning tokens level. | o1 | o1 | 142.74 | 140.13 | 147.88 | |||
| o1-pro | o1-pro-2025-03-19_low | o1 Pro (low) | OpenAI | United States of America | 2025-03-19 | API access | 2025-03-19 | Likely | An An enhanced version of OpenAI's first reasoning model,o1, evaluated with low reasoning effort. | o1-pro | o1-pro | ||||||
| ml-elephant | Zoo Ml-Ephant | Unverified | ml-elephant | ||||||||||||||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_15K | Claude 3.7 Sonnet (15k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 15,000 reasoning tokens allowed. | Claude 3.7 Sonnet (15k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 141.83 | 139.34 | 147.11 | |
| Pixtral 12B | Pixtral-12B-2409 | Mistral AI | France | 2024-09-17 | Open weights (unrestricted) | 2024-09-17 | Confident | Pixtral 12B | pixtral-12b | ||||||||
| Claude Opus 4.5 | claude-opus-4-5-20251101_128K | Claude Opus 4.5 (128k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (128k thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.62 | 146.60 | 155.89 | |||
| gpt-oss-120b | gpt-oss-120b_unknown | gpt-oss-120b (unknown thinking) | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | The larger one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-120b | gpt-oss-120b | 140.55 | 135.47 | 146.61 | ||
| Qwen3.5-9B | qwen3.5-9B | Alibaba | China | 2026-02-24 | Open weights (restricted use) | 2026-02-24 | Confident | Qwen3.5-9B | qwen3-5-9b | ||||||||
| gpt-oss-20b | gpt-oss-20b_unknown | gpt-oss-20b (unknown thinking) | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | Confident | The smaller one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-20b | gpt-oss-20b | |||||
| DeepSeek-V3.2-Exp | DeepSeek-V3.2-Exp_unknown | DeepSeek | China | 2025-09-29 | Open weights (unrestricted) | 2025-09-29 | 4.2e+24 | Confident | DeepSeek-V3.2-Exp | deepseek-v3-2-exp | 144.94 | 142.23 | 150.68 | ||||
| DeepSeek-R1 (May 2025) | chutes/DeepSeek-R1-0528 | DeepSeek-R1 (May 2025) | DeepSeek | China | 2025-05-28 | Open weights (unrestricted) | 2025-05-28 | 4.0e+24 | Confident | DeepSeek-R1 (May 2025) (Chutes) | DeepSeek-R1 (May 2025) | deepseek-r1-may-2025 | 142.00 | 139.40 | 147.19 | ||
| chutes/GLM-4.5-FP8 | 2025-07-27 | Unverified | chutes-glm-4-5-fp8 | ||||||||||||||
| gpt-oss-120b | chutes/gpt-oss-120b | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | gpt-oss-120b (Chutes) | gpt-oss-120b | gpt-oss-120b | 140.55 | 135.47 | 146.61 | |||
| gpt-oss-120b | chutes/gpt-oss-120b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | gpt-oss-120b (high) (Chutes) | gpt-oss-120b | gpt-oss-120b | 140.55 | 135.47 | 146.61 | |||
| Qwen3-8B | chutes/Qwen3-8B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 1.8e+24 | Confident | Qwen3 8B (Chutes) | Qwen3-8B | qwen3-8b | ||||||
| Qwen3-235B-A22B | chutes/Qwen3-235B-A22B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Likely | Qwen3 (Apr 2025) (Chutes) | Qwen3-235B-A22B | qwen3-235b-a22b | 139.61 | 134.81 | 144.40 | |||
| Qwen3-235B-A22B (Jul 2025) | chutes/Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | Qwen3 (Jul 2025) (Chutes) | Qwen3-235B-A22B-Thinking (Jul 2025) | Qwen3-235B-A22B-Thinking (Jul 2025) | qwen3-235b-a22b-thinking-jul-2025 | 144.95 | 141.90 | 151.93 | ||
| QwQ-32B | chutes/QwQ-32B | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B (Chutes) | QwQ-32B | QwQ-32B | qwq-32b | |||||
| deepinfra/Qwen3-Next-80B-A3B-Instruct | 2025-09-11 | Unverified | deepinfra-qwen3-next-80b-a3b-instruct | ||||||||||||||
| DeepSeek-V3.2-Exp | DeepSeek-V3.2-Exp_high | DeepSeek | China | 2025-09-29 | Open weights (unrestricted) | 2025-09-29 | 4.2e+24 | Confident | DeepSeek-V3.2-Exp | deepseek-v3-2-exp | 144.94 | 142.23 | 150.68 | ||||
| Gemini 2.5 Flash-Lite | gemini-2.5-flash-lite-preview-06-17-thinking | gemini-2.5-flash-lite-preview-06-17 (Default thinking length) | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-15 | Unknown | A June 2025 preview of Gemini 2.5 Flash Lite. | Gemini 2.5 Flash-Lite (Jun 2025) | Gemini 2.5 Flash-Lite (Jun 2025) | Gemini 2.5 Flash-Lite (Jun 2025) | gemini-2-5-flash-lite-jun-2025 | ||||
| DeepSeek-V3 (Mar 2025) | chutes/DeepSeek-V3-0324 | DeepSeek-V3 (Mar 2025) | DeepSeek | China | 2025-03-24 | Open weights (restricted use) | 2025-03-24 | 3.3e+24 | Confident | DeepSeek-V3 (Mar 2025) (Chutes) | DeepSeek-V3 (Mar 2025) | deepseek-v3-mar-2025 | 137.07 | 134.38 | 142.52 | ||
| Gemma 3 27B | chutes/Gemma-3-27b-It | Google DeepMind | United States of America | 2025-03-11 | Open weights (restricted use) | 2025-03-12 | 2.3e+24 | Confident | Gemma-3-27b-it (Chutes) | Gemma 3 27B | gemma-3-27b | 131.05 | 127.90 | 135.66 | |||
| Kimi K2 | fireworks/Kimi-K2-Instruct-0905 | Kimi K2 Instruct (0905) | Moonshot | China | 2025-09-05 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | Kimi K2 (Sep 2025) | Kimi K2 (Sep 2025) | kimi-k2-sep-2025 | 141.04 | 137.75 | 147.36 | ||
| Grok-3 mini | grok-3-mini-beta_medium | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | Unknown | A beta release of a smaller version of XAI's third generation Grok model, evaluated with medium reasoning effort. | Grok-3 mini | grok-3-mini | 141.05 | 138.50 | 145.60 | ||||
| MiniMax-M1-80k | MiniMax-M1-80k | MiniMax | China | 2025-06-13 | Open weights (unrestricted) | 2025-06-13 | 4.3e+24 | Likely | MiniMax-M1-80k | minimax-m1-80k | |||||||
| nvidia-nemotron-nano-9b-v2 | 2025-08-18 | Unverified | nvidia-nemotron-nano-9b-v2 | ||||||||||||||
| Qwen3-235B-A22B (Jul 2025) | parasail-qwen3-235b-a22b-instruct-2507 | Alibaba | China | 2025-09-01 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | Qwen3 (Parasail) | Qwen3-235B-A22B-Instruct (Jul 2025) | Qwen3-235B-A22B-Instruct (Jul 2025) | qwen3-235b-a22b-instruct-jul-2025 | 138.99 | 135.96 | 144.81 | ||
| Qwen3-30B-A3B | chutes/Qwen3-30B-A3B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 6.5e+23 | Likely | Qwen3-30B-A3B (Chutes) | Qwen3-30B-A3B | qwen3-30b-a3b | ||||||
| Qwen3-14B | chutes/Qwen3-14B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 3.2e+24 | Confident | Qwen3 14B | Qwen3-14B | qwen3-14b | ||||||
| Qwen3-32B | chutes/Qwen3-32B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 7.1e+24 | Confident | Qwen3 32B (Chutes) | Qwen3-32B | qwen3-32b | ||||||
| chutes/Qwen3-Next-80B-A3B-Instruct | 2025-09-11 | Unverified | chutes-qwen3-next-80b-a3b-instruct | ||||||||||||||
| Llama 4 Scout | chutes/Llama-4-Scout-17B-16E Instruct | Meta AI | United States of America | 2025-04-06 | Open weights (restricted use) | 2025-04-05 | 4.1e+24 | Likely | Llama 4 Maverick (16 experts) (Chutes) | Llama 4 Scout | llama-4-scout | 130.52 | 127.97 | 134.52 | |||
| claude-mythos-preview-early | Unverified | claude-mythos-preview-early | |||||||||||||||
| GPT-3.5 Turbo Instruct | gpt-3.5-turbo-instruct | OpenAI | United States of America | 2023-09-18 | API access | 2023-09-28 | Likely | An instruction-tuned version of OpenAI's GPT-3.5 turbo. | GPT-3.5 Turbo Instruct | gpt-3-5-turbo-instruct | |||||||
| GPT-3 175B (davinci) | davinci-002 | OpenAI | United States of America | API access | 2020-05-28 | 3.1e+23 | Confident | GPT-3 175B (davinci-002) | GPT-3 175B (davinci) | gpt-3-175b-davinci | |||||||
| GPT-5.4 Pro | gpt-5.4-pro-2026-03-05_none | GPT-5.4 Pro (no thinking) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 Pro (no thinking) | GPT-5.4 Pro | gpt-5-4-pro | 157.91 | 155.37 | 165.49 | |||
| GPT‑5-Codex | gpt-5-codex_high | GPT-5-codex (high) | OpenAI | United States of America | 2025-09-15 | API access | 2025-09-15 | Unknown | GPT-5-codex | GPT‑5-Codex | gpt5-codex | ||||||
| Grok-3 mini | grok-3-mini_high | Grok 3 mini | xAI | United States of America | 2025-06-24 | API access | 2025-02-19 | Unknown | Grok 3 mini (high) | Grok-3 mini | grok-3-mini | 141.05 | 138.50 | 145.60 | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-09-2025_16k | Google DeepMind | United States of America | 2025-09-25 | API access | 2025-04-17 | Unknown | Gemini 2.5 Flash-Lite (Sep 2025, 16k thinking) | Gemini 2.5 Flash (Sep 2025) | Gemini 2.5 Flash (Sep 2025) | gemini-2-5-flash-sep-2025 | 143.18 | 134.59 | 150.09 | |||
| gpt-oss-120b | gpt-oss-120b_medium | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | The larger one of OpenAI's two latest open source language models, evaluated with medium reasoning effort. | gpt-oss-120b | gpt-oss-120b | 140.55 | 135.47 | 146.61 | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-04-17 (16K thinking) | Google DeepMind | United States of America | 2025-04-17 | API access | 2025-04-17 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 16,000 reasoning tokens. | Gemini 2.5 Flash (Apr 2025) | Gemini 2.5 Flash (Apr 2025) | gemini-2-5-flash-apr-2025 | 140.55 | 136.89 | 146.50 | |||
| GLM-4.5 | glm-4.5_thinking | GLM-4.5 Thinking | Z.ai (Zhipu AI),Tsinghua University | China | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.4e+24 | Confident | GLM-4.5 | glm-4-5 | ||||||
| GPT-5 | gpt-5-chat | OpenAI | United States of America | API access | 2025-08-07 | 6.6e+25 | Speculative | GPT-5 | gpt-5 | 150.00 | 147.33 | 156.34 | |||||
| Kimi K2 | kimi-k2-0711-preview | Moonshot | China | 2025-07-12 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | A preview of an updated version of Kimi K2. | Kimi K2 (Jul 2025) | Kimi K2 (Jul 2025) | kimi-k2-jul-2025 | 140.42 | 137.12 | 145.88 | ||
| Qwen3-235B-A22B (Jul 2025) | qwen/qwen3-235b-a22b-thinking-2507 | Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | Qwen3-235B-A22B-Thinking (Jul 2025) | Qwen3-235B-A22B-Thinking (Jul 2025) | qwen3-235b-a22b-thinking-jul-2025 | 144.95 | 141.90 | 151.93 | ||
| DeepSeek-V3.1 | DeepSeek-V3.1_thinking | DeepSeek | China | 2025-08-21 | Open weights (unrestricted) | 2025-08-21 | 3.6e+24 | Confident | DeepSeek-V3.1 (thinking) | DeepSeek-V3.1 | deepseek-v3-1 | 138.75 | 135.48 | 145.60 | |||
| GPT-5.4 Nano | gpt-5.4-nano-2026-03-17_none | GPT-5.4 nano (no thinking) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 nano (no thinking) | GPT-5.4 Nano | gpt-5-4-nano | 146.54 | 143.59 | 152.29 | |||
| GPT-5.4 Mini | gpt-5.4-mini-2026-03-17_none | GPT-5.4 mini (none) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 mini (none) | GPT-5.4 Mini | gpt-5-4-mini | 148.96 | 145.80 | 155.00 | |||
| gpt-oss-20b | gpt-oss-20b_medium | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | Confident | The smaller one of OpenAI's two latest open source language models, evaluated with medium reasoning effort. | gpt-oss-20b | gpt-oss-20b | ||||||
| Gemini 2.5 Flash-Lite | gemini-2.5-flash-lite-preview-06-17_16K | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-15 | Unknown | A June 2025 preview of Gemini 2.5 Flash Lite, evaluated with up to 16,000 reasoning tokens. | Gemini 2.5 Flash-Lite (Jun 2025, 16k thinking) | Gemini 2.5 Flash-Lite (Jun 2025) | Gemini 2.5 Flash-Lite (Jun 2025) | gemini-2-5-flash-lite-jun-2025 | |||||
| qwen3-coder-next | 2026-02-02 | Unverified | qwen3-coder-next | ||||||||||||||
| mistral-medium-2508 | 2025-08-12 | Unverified | mistral-medium-2508 | ||||||||||||||
| Qwen3-30B-A3B | Qwen3-30B-A3B | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 6.5e+23 | Likely | Qwen3-30B-A3B | Qwen3-30B-A3B | qwen3-30b-a3b | ||||||
| QwQ-32B | QwQ-32B (16K thinking) | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B (16k thinking) | QwQ-32B | QwQ-32B | qwq-32b | |||||
| GPT-5 Pro | gpt-5-pro-2025-10-06_medium | GPT-5 Pro | OpenAI | United States of America | 2025-10-07 | API access | 2025-10-07 | Unknown | GPT-5 Pro (medium) | GPT-5 Pro | gpt-5-pro | 150.48 | 147.41 | 157.34 | |||
| gemini-robotics-er-1.5-preview | 2025-09-26 | Unverified | gemini-robotics-er-1-5-preview | ||||||||||||||
| Gemini 2.5 Flash-Lite | gemini-2.5-flash-lite-preview-09-2025 | Google DeepMind | United States of America | 2025-09-25 | API access | 2025-06-15 | Unknown | Gemini 2.5 Flash-Lite (Sep 2025) | Gemini 2.5 Flash-Lite (Sep 2025) | Gemini 2.5 Flash-Lite (Sep 2025) | gemini-2-5-flash-lite-sep-2025 | ||||||
| computer-use-preview-2025-03-11 | 2025-03-11 | Unverified | computer-use-preview-2025-03-11 |
The chart above shows the ECI frontier over time, where each step marks a new model taking the top spot. Claude 3 Opus overtook GPT-4 in February 2024, the first of 17 leading models through June 2026, with each holding the lead for a median of about seven weeks. We begin the comparison at GPT-4’s release date, March 14, 2023.
Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.
Learn more about this graph
We track which model has the highest ECI at each point in time, and how long it stays at the top of the leaderboard. The ECI can fluctuate slightly over time, though rarely enough to significantly shuffle model rankings. That said, this is a retrospective view and does not account for changes to ECI estimates since models’ release dates.
Data
The ECI is computed from over 50 benchmarks on Epoch’s Benchmarking Hub, which combines Epoch’s internal evaluations with scores reported by model developers and benchmark creators. The index currently covers 168 models released since January 1, 2023. We restrict attention to models on the ECI frontier, meaning those that surpassed every earlier model on release. We limit analysis for this Data Insight to models released since March 14, 2023, because earlier models have especially sparse data and correspondingly noisy ECI estimates.
Analysis
For each frontier model, we measure the number of days between its release and the release of the next frontier model. The current leader, GPT-5.5 Pro, has a lead measured only up to the latest available data.
GPT-4 led for 352 days, from its release on March 14, 2023 until Claude 3 Opus surpassed it on February 29, 2024. This is by far the longest lead in the data. The next-longest belongs to OpenAI’s o1, at 98 days, making GPT-4’s lead about 3.6 times as long. Among the 17 models that have led since GPT-4, none has held the top for a three-digit number of days, and the median lead has been about seven weeks.
| Model | Reached the top | Days at the top |
|---|---|---|
| GPT-4 | Mar 14, 2023 | 352 |
| o1 | Dec 17, 2024 | 98 |
| o1-mini | Sep 12, 2024 | 96 |
| Claude 3.5 Sonnet | Jun 20, 2024 | 84 |
| GPT-5 | Aug 7, 2025 | 61 |
| o3-pro | Jun 10, 2025 | 58 |
Assumptions and limitations
- No pre-2023 coverage: Our public ECI does not score pre-2023 models, so we can only compare GPT-4 to models released after it. Models released before 2023 may also have led for extended periods, though we estimate this is unlikely, based on provisional high-uncertainty ECI scores.
- No significance threshold: We rank models by their point-estimate ECI and treat any increase as a change at the top. Some leads since GPT-4 end on small ECI differences that fall within confidence intervals. If we considered only statistically significant changes, several lead periods would lengthen moderately.
- ECI is periodically refit: As benchmarks and models are added, historical scores can shift slightly between updates.
- The current lead: The latest lead is measured only up to the most recent data, meaning it could extend beyond its current length.

