Since January 2026, the most capable open-weight models have lagged frontier closed models by an average of four months in the Epoch Capabilities Index (ECI), our aggregate measure of model capability. The average ECI gap was 8 points, similar to the gap between GPT-5 and GPT-5.5.
| Model | Model version ID | Display name | Organization | Country (of organization) | Version release date | Model accessibility | Publication date | Training compute (FLOP) | Confidence | Description | Unique display name | Model aggregation | Model group | Slug | ECI | ECI CI low | ECI CI high |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemini 3.5 Flash | gemini-3.5-flash_high | Gemini 3.5 Flash (high) | United States of America | 2026-05-19 | Unverified | Gemini 3.5 Flash | gemini-3-5-flash | 156.31 | 154.32 | 164.67 | |||||||
| Qwen 3.6 Max (Preview) | qwen3.6-max-preview | Qwen 3.6 Max (Preview) | Alibaba | China | 2026-04-20 | Unverified | Qwen 3.6 Max (Preview) | qwen-3-6-max-preview | 150.16 | 147.01 | 159.05 | ||||||
| GLM-5.1 | glm-5.1 | GLM-5.1 | Z.ai (Zhipu AI) | China | 2026-04-07 | 2026-04-07 | Confident | GLM-5.1 | glm-5-1 | 149.94 | 147.55 | 157.07 | |||||
| Qwen 3.5 Plus (hosted 397B-A17B) | qwen3.5-plus | Alibaba | China | 2026-02-16 | 2026-02-16 | Likely | Qwen 3.5 Plus (hosted 397B-A17B) | qwen-3-5-plus-hosted-397b-a17b | 146.11 | 143.00 | 152.98 | ||||||
| Qwen 3.6 Plus | qwen3.6-plus | Qwen 3.6 Plus (2026-04-02) | Alibaba | China | 2026-03-31 | API access | 2026-04-01 | Confident | Qwen 3.6 Plus | qwen-3-6-plus | 149.08 | 145.48 | 156.15 | ||||
| Qwen 3.6 Flash | qwen3.6-flash | Qwen 3.6 Flash (2026-04-16) | 2026-04-27 | Unverified | Qwen 3.6 Flash | qwen-3-6-flash | 145.16 | 142.46 | 152.12 | ||||||||
| Qwen 3.5 Flash (hosted 35B-A3B) | qwen3.5-flash | Alibaba | China | 2026-02-25 | 2026-02-25 | Likely | Qwen 3.5 Flash (hosted 35B-A3B) | qwen-3-5-flash-hosted-35b-a3b | 144.88 | 140.98 | 151.42 | ||||||
| AI Co-Mathematician | gdm-ai-co-mathematician | AI co-mathematician | Google DeepMind | United States of America | 2026-05-08 | Unreleased | 2026-05-08 | Unverified | AI Co-Mathematician | ai-co-mathematician | |||||||
| Kimi K2.6 | kimi-k2.6 | Kimi K2.6 | Moonshot | China | 2026-04-20 | Open weights (unrestricted) | 2026-04-20 | Confident | Kimi K2.6 | kimi-k2-6 | 151.60 | 148.44 | 159.28 | ||||
| GPT-5.5 | gpt-5.5_low | GPT-5.5 (low) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.22 | 155.43 | 165.84 | ||||
| GPT-5.5 Pro | gpt-5.5-pro-pre-release_xhigh | GPT-5.5 Pro (xhigh) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 Pro | gpt-5-5-pro | 159.35 | 155.79 | 167.47 | ||||
| GPT-5.5 | gpt-5.5-pre-release_xhigh | GPT-5.5 (xhigh) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.22 | 155.43 | 165.84 | ||||
| GPT-5.5 Pro | gpt-5.5-pro-pre-release_high | GPT-5.5 Pro (high) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 Pro | gpt-5-5-pro | 159.35 | 155.79 | 167.47 | ||||
| Claude Opus 4.7 | claude-opus-4-7_xhigh | Claude Opus 4.7 (xhigh) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (xhigh) | Claude Opus 4.7 | claude-opus-4-7 | 156.18 | 153.09 | 164.09 | |||
| Claude Opus 4.7 | claude-opus-4-7_max | Claude Opus 4.7 (max) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (max) | Claude Opus 4.7 | claude-opus-4-7 | 156.18 | 153.09 | 164.09 | |||
| GPT-5.4 Mini | gpt-5.4-mini-2026-03-17_high | GPT-5.4 mini (high) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 mini (high) | GPT-5.4 Mini | gpt-5-4-mini | 148.74 | 145.69 | 155.92 | |||
| GPT-5.4 Nano | gpt-5.4-nano-2026-03-17_high | GPT-5.4 nano (high) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 nano (high) | GPT-5.4 Nano | gpt-5-4-nano | 146.23 | 142.65 | 153.80 | |||
| Muse Spark | muse-spark | Muse Spark | Meta AI | United States of America | 2026-04-08 | API access | 2026-04-08 | Unverified | Muse Spark | Muse Spark | muse-spark | 155.11 | 151.48 | 163.68 | |||
| Gemma 4 31B IT | gemma-4-31b-it | Gemma 4 31B IT | Google DeepMind | United States of America | 2026-04-02 | Open weights (restricted use) | 2026-04-02 | Likely | Gemma 4 31B IT | Gemma 4 31B IT | gemma-4-31b-it | ||||||
| GPT-5.4 | gpt-5.4-2026-03-05_high | GPT-5.4 (high) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (high) | GPT-5.4 | gpt-5-4 | 156.13 | 153.44 | 163.96 | |||
| GPT-5.4 | gpt-5.4-2026-03-05_xhigh | GPT-5.4 (xhigh) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (xhigh) | GPT-5.4 | gpt-5-4 | 156.13 | 153.44 | 163.96 | |||
| GPT-5.4 Pro | gpt-5.4-pro-2026-03-05_xhigh | GPT-5.4 Pro (xhigh) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 Pro (xhigh) | GPT-5.4 Pro | gpt-5-4-pro | 157.73 | 155.08 | 165.39 | |||
| GPT-5.4 Pro | gpt-5.4-pro-2026-03-05-web-app | GPT-5.4 Pro (web) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 Pro (Web App) | GPT-5.4 Pro | gpt-5-4-pro | 157.73 | 155.08 | 165.39 | |||
| GPT-5.3 Codex | gpt-5.3-codex_high | GPT-5.3 Codex (high) | OpenAI | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | GPT-5.3 Codex | gpt-5-3-codex | 155.89 | 153.05 | 164.24 | ||||
| Gemini 3.1 Pro | gemini-3.1-pro-preview-customtools | Gemini 3.1 Pro Preview | Google DeepMind | United States of America | 2026-02-19 | API access | 2026-02-19 | Likely | Gemini 3.1 Pro | gemini-3-1-pro | 156.63 | 154.43 | 164.92 | ||||
| Claude Sonnet 4.6 | claude-sonnet-4-6_16K | Claude Sonnet 4.6 (16k thinking) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 152.58 | 148.95 | 158.53 | ||||
| Claude Sonnet 4.6 | claude-sonnet-4-6 | Claude Sonnet 4.6 | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 152.58 | 148.95 | 158.53 | ||||
| Claude Sonnet 4.6 | claude-sonnet-4-6_32K | Claude Sonnet 4.6 (32k thinking) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 152.58 | 148.95 | 158.53 | ||||
| Gemini 3.1 Pro | gemini-3.1-pro-preview | Gemini 3.1 Pro Preview | Google DeepMind | United States of America | 2026-02-19 | API access | 2026-02-19 | Likely | Gemini 3.1 Pro | gemini-3-1-pro | 156.63 | 154.43 | 164.92 | ||||
| Claude Opus 4.6 | claude-opus-4-6_120K | Claude Opus 4.6 (120k thinking) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (120k thinking) | Claude Opus 4.6 | claude-opus-4-6 | 155.28 | 152.40 | 163.08 | |||
| GLM-5 | glm-5 | GLM-5 | Z.ai (Zhipu AI) | China | 2026-02-11 | Open weights (unrestricted) | 2026-02-11 | Likely | GLM-5 | glm-5 | 146.62 | 143.90 | 152.95 | ||||
| Gemini 3 Flash | gemini-3-flash-preview | Google DeepMind | United States of America | 2025-12-17 | API access | 2025-12-17 | Unknown | Gemini 3 Flash | gemini-3-flash | 150.94 | 143.97 | 157.85 | |||||
| Claude Opus 4.6 | claude-opus-4-6 | Claude Opus 4.6 (no thinking) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (no thinking) | Claude Opus 4.6 | claude-opus-4-6 | 155.28 | 152.40 | 163.08 | |||
| GPT-5.1 | gpt-5.1-2025-11-13_high | GPT-5.1 (high) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (high) | GPT-5.1 | gpt-5-1 | 149.71 | 146.98 | 156.51 | |||
| Kimi K2.5 | kimi-k2.5 | Kimi K2.5 | Moonshot | China | 2026-01-27 | Open weights (unrestricted) | 2026-02-02 | 5.8e+24 | Likely | Kimi K2.5 | kimi-k2-5 | 148.22 | 145.05 | 154.59 | |||
| Claude Opus 4.6 | claude-opus-4-6_max | Claude Opus 4.6 (max) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (max) | Claude Opus 4.6 | claude-opus-4-6 | 155.28 | 152.40 | 163.08 | |||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-05 | Unknown | Google’s general-purpose frontier reasoning model. | Gemini 2.5 Pro | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.72 | 143.80 | 153.41 | ||
| Gemini 3 Pro | gemini-3-pro-preview | Gemini 3 Pro Preview | Google DeepMind | United States of America | 2025-11-18 | API access | 2025-11-18 | Unknown | Gemini 3 Pro Preview | Gemini 3 Pro | gemini-3-pro | 153.45 | 150.79 | 161.36 | |||
| GPT-5.2 | gpt-5.2-2025-12-11_high | GPT-5.2 (high) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (high) | GPT-5.2 | gpt-5-2 | 153.72 | 150.93 | 161.15 | |||
| o3 | o3-2025-04-16_medium | o3 (medium) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | Unknown | OpenAI's second generation o-series reasoning model, evaluated with medium reasoning effort. | o3 | o3 | 147.27 | 144.19 | 154.26 | |||
| Claude Opus 4.1 | claude-opus-4-1-20250805 | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model. | Claude Opus 4.1 | Claude Opus 4.1 | claude-opus-4-1 | 144.73 | 141.67 | 150.99 | |||
| GPT-4o | gpt-4o-2024-11-20 | GPT-4o (Nov 2024) | OpenAI | United States of America | 2024-11-20 | API access | 2024-05-13 | Speculative | The November 2024 version of GPT-4o, OpenAI's then-flagship multimodal language model. | GPT-4o (Nov 2024) | GPT-4o (Nov 2024) | gpt-4o-nov-2024 | 129.36 | 127.15 | 134.73 | ||
| GPT-4.1 | gpt-4.1-2025-04-14 | GPT-4.1 | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | Unknown | A coding-optimized GPT-4 series model from OpenAI. | GPT-4.1 | gpt-4-1 | 137.60 | 135.09 | 142.94 | |||
| Claude Opus 4.6 | claude-opus-4-6_64K | Claude Opus 4.6 (64k thinking) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (64k thinking) | Claude Opus 4.6 | claude-opus-4-6 | 155.28 | 152.40 | 163.08 | |||
| Claude Opus 4.6 | claude-opus-4-6_32K | Claude Opus 4.6 (32k thinking) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (32k thinking) | Claude Opus 4.6 | claude-opus-4-6 | 155.28 | 152.40 | 163.08 | |||
| GPT-5 | gpt-5-2025-08-07_high | GPT-5 (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 (high) | GPT-5 | gpt-5 | 150.00 | 146.38 | 157.62 | |
| Claude Opus 4 | claude-opus-4-20250514 | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1). | Claude Opus 4 | Claude Opus 4 | claude-opus-4 | 143.37 | 140.42 | 150.37 | ||
| Claude Opus 4.5 | claude-opus-4-5-20251101 | Claude Opus 4.5 (no thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (no thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.89 | 147.32 | 157.25 | |||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929 | Claude Sonnet 4.5 (no thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Anthropic's latest mid-size model, part of the Claude 4.5 series. | Claude Sonnet 4.5 (no thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 147.15 | 144.42 | 153.64 | ||
| GPT-5 | gpt-5-2025-08-07_medium | GPT-5 (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 (medium) | GPT-5 | gpt-5 | 150.00 | 146.38 | 157.62 | |
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219 | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic. | Claude 3.7 Sonnet | Claude 3.7 Sonnet | claude-3-7-sonnet | 142.03 | 139.52 | 148.45 | ||
| Kimi K2.5 | fireworks/kimi-k2p5 | Kimi K2.5 (Fireworks) | Moonshot | China | 2026-01-27 | Open weights (unrestricted) | 2026-02-02 | 5.8e+24 | Likely | Kimi K2.5 | kimi-k2-5 | 148.22 | 145.05 | 154.59 | |||
| GPT-5 mini | gpt-5-mini-2025-08-07_medium | GPT-5 mini (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 mini (medium) | GPT-5 mini | gpt-5-mini | 145.59 | 142.41 | 152.17 | ||
| Grok 4 | grok-4-0709 | xAI | United States of America | 2025-07-09 | API access | 2025-07-09 | 5.0e+26 | Speculative | Grok 4 | Grok 4 | grok-4 | 147.43 | 144.38 | 154.25 | |||
| GLM-4.7 | zai-org/GLM-4.7 | GLM-4.7 (Together) | Z.ai (Zhipu AI) | China | 2025-12-22 | Open weights (unrestricted) | 2025-12-22 | 4.4e+24 | Likely | GLM-4.7 | glm-4-7 | 144.61 | 141.55 | 150.45 | |||
| GLM-4.7 | glm-4.7 | GLM-4.7 | Z.ai (Zhipu AI) | China | 2025-12-22 | Open weights (unrestricted) | 2025-12-22 | 4.4e+24 | Likely | GLM-4.7 | GLM-4.7 | glm-4-7 | 144.61 | 141.55 | 150.45 | ||
| GPT-5.2 Pro | gpt-5.2-pro-2025-12-11-webapp | GPT-5.2 Pro (web) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Likely | GPT-5.2 Pro (Web App) | GPT-5.2 Pro | gpt-5-2-pro | 154.19 | 150.15 | 162.75 | |||
| DeepSeek-V3.2 | fireworks/deepseek-v3p2 | DeepSeek-V3.2 (Thinking; Fireworks) | DeepSeek | China | 2025-12-01 | Open weights (unrestricted) | 2025-12-01 | 4.2e+24 | Likely | DeepSeek-V3.2 | deepseek-v3-2 | 146.46 | 143.27 | 152.46 | |||
| Gemini 2.5 Flash | gemini-2.5-flash | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-04-17 | Unknown | A small model from Google DeepMind's Gemini 2.5 series. | Gemini 2.5 Flash (Jun 2025) | Gemini 2.5 Flash (Jun 2025) | Gemini 2.5 Flash (Jun 2025) | gemini-2-5-flash-jun-2025 | 140.38 | 135.93 | 146.01 | ||
| DeepSeek-V3.2 | deepseek-reasoner | DeepSeek-V3.2 (Thinking) | DeepSeek | China | 2025-12-01 | Open weights (unrestricted) | 2025-12-01 | 4.2e+24 | Likely | DeepSeek-V3.2 | deepseek-v3-2 | 146.46 | 143.27 | 152.46 | |||
| gpt-oss-120b | openai/gpt-oss-120b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | gpt-oss-120b (high) | gpt-oss-120b | gpt-oss-120b | 140.79 | 135.32 | 148.00 | |||
| GPT-5.2 | gpt-5.2-2025-12-11_xhigh | GPT-5.2 (xhigh) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (xhigh) | GPT-5.2 | gpt-5-2 | 153.72 | 150.93 | 161.15 | |||
| Qwen3-235B-A22B (Jul 2025) | Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | A 22 billion active (235 billion total)-parameter mixture-of-experts reasoning model in the Qwen 3 series. | Qwen3-235B-A22B-Thinking (Jul 2025) | Qwen3-235B-A22B-Thinking (Jul 2025) | qwen3-235b-a22b-thinking-jul-2025 | 145.81 | 142.32 | 152.99 | ||
| GPT-5.2 | gpt-5.2-2025-12-11_medium | GPT-5.2 (medium) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (medium) | GPT-5.2 | gpt-5-2 | 153.72 | 150.93 | 161.15 | |||
| GPT-5.2 | gpt-5.2-2025-12-11_low | GPT-5.2 (low) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (low) | GPT-5.2 | gpt-5-2 | 153.72 | 150.93 | 161.15 | |||
| Qwen3-235B-A22B (Jul 2025) | Qwen/Qwen3-235B-A22B-Thinking-2507 | Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | Qwen3-235B-A22B-Thinking (Jul 2025) | Qwen3-235B-A22B-Thinking (Jul 2025) | qwen3-235b-a22b-thinking-jul-2025 | 145.81 | 142.32 | 152.99 | ||
| Qwen3-Max | qwen3-max-2025-09-23 | Qwen3-Max-Instruct | Alibaba | China | 2025-09-24 | API access | 2025-09-05 | 1.5e+25 | Speculative | A 1-trillion total parameter scale model in the Qwen 3 series. | Qwen3 Max | Qwen3-Max | qwen3-max | 144.94 | 139.82 | 159.43 | |
| Kimi K2 Thinking | kimi-k2-thinking-turbo | Kimi K2 Thinking Turbo | Moonshot | China | 2025-11-06 | Open weights (restricted use) | 2025-11-06 | 4.2e+24 | Likely | Kimi K2 Thinking | kimi-k2-thinking | 145.60 | 142.44 | 152.09 | |||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_59K | Claude Sonnet 4.5 (59k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (59k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 147.15 | 144.42 | 153.64 | |||
| Claude Opus 4.1 | claude-opus-4-1-20250805_27K | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model, benchmarked with up to 27,000 reasoning tokens. | Claude Opus 4.1 (27k thinking) | Claude Opus 4.1 | claude-opus-4-1 | 144.73 | 141.67 | 150.99 | |||
| GPT-5 mini | gpt-5-mini-2025-08-07_high | GPT-5 mini (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 mini (high) | GPT-5 mini | gpt-5-mini | 145.59 | 142.41 | 152.17 | ||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_32K | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (32k thinking) | Claude Haiku 4.5 | claude-haiku-4-5 | 143.09 | 140.17 | 149.33 | ||||
| o4-mini | o4-mini-2025-04-16_high | o4-mini (high) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | Unknown | The latest smaller size o-series reasoning model from OpenAI, evaluated with high reasoning effort. | o4-mini | o4-mini | 146.91 | 143.73 | 153.62 | |||
| Claude Opus 4.5 | claude-opus-4-5-20251101_32K | Claude Opus 4.5 (32k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (32k thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.89 | 147.32 | 157.25 | |||
| o3 | o3-2025-04-16_high | o3 (high) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | Unknown | OpenAI's second generation o-series reasoning model, evaluated with high reasoning effort. | o3 | o3 | 147.27 | 144.19 | 154.26 | |||
| GPT-5 nano | gpt-5-nano-2025-08-07_high | GPT-5 nano (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 nano (high) | GPT-5 nano | gpt-5-nano | 140.95 | 137.13 | 147.50 | ||
| Grok-3 mini | grok-3-mini-beta_high | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | Unknown | A beta release of a smaller version of XAI's third generation Grok model, evaluated with high reasoning effort. | Grok-3 mini | grok-3-mini | 141.19 | 138.81 | 147.60 | ||||
| Kimi K2 Thinking | moonshotai/Kimi-K2-Thinking | Kimi K2 Thinking (Together) | Moonshot | China | 2025-11-06 | Open weights (restricted use) | 2025-11-06 | 4.2e+24 | Likely | Kimi K2 Thinking | kimi-k2-thinking | 145.60 | 142.44 | 152.09 | |||
| GLM-4.6 | zai-org/GLM-4.6 | GLM-4.6 (Together) | Z.ai (Zhipu AI),Tsinghua University | China | 2025-09-30 | Open weights (unrestricted) | 2025-09-30 | 4.4e+24 | Likely | GLM-4.6 | glm-4-6 | 141.38 | 132.95 | 147.00 | |||
| DeepSeek-R1 (May 2025) | DeepSeek-R1-0528 | DeepSeek-R1 (May 2025) | DeepSeek | China | 2025-05-28 | Open weights (unrestricted) | 2025-05-28 | 4.0e+24 | Confident | An updated May 2025 version of DeepSeek's reasoning model, R1. | DeepSeek-R1 (May 2025) | DeepSeek-R1 (May 2025) | deepseek-r1-may-2025 | 142.19 | 139.19 | 148.16 | |
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_32K | Claude Sonnet 4.5 (32k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Anthropic's latest mid-size model, part of the Claude 4.5 series, evaluated with up to 32,000 reasoning tokens. | Claude Sonnet 4.5 (32k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 147.15 | 144.42 | 153.64 | ||
| o3-mini | o3-mini-2025-01-31_high | o3-mini (high) | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | Unknown | A smaller version of OpenAI's o3 reasoning model, evaluated with high reasoning effort. | o3-mini | o3-mini | 141.63 | 138.79 | 147.93 | |||
| Claude 3.5 Haiku | claude-3-5-haiku-20241022 | Claude 3.5 Haiku (Oct 2024) | Anthropic | United States of America | 2024-10-22 | API access | 2024-10-22 | Unknown | The smallest model in Anthropic’s Claude 3.5 Series. | Claude 3.5 Haiku (Oct 2024) | Claude 3.5 Haiku | claude-3-5-haiku | 127.57 | 121.86 | 133.01 | ||
| GPT-5.1 | gpt-5.1-2025-11-13_none | GPT-5.1 (no thinking) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (no thinking) | GPT-5.1 | gpt-5-1 | 149.71 | 146.98 | 156.51 | |||
| GPT-5.1 | gpt-5.1-2025-11-13_low | GPT-5.1 (low) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (low) | GPT-5.1 | gpt-5-1 | 149.71 | 146.98 | 156.51 | |||
| Claude Opus 4.5 | claude-opus-4-5-20251101_16K | Claude Opus 4.5 (16k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (16k thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.89 | 147.32 | 157.25 | |||
| GPT-5.1 | gpt-5.1-2025-11-13_medium | GPT-5.1 (medium) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (medium) | GPT-5.1 | gpt-5-1 | 149.71 | 146.98 | 156.51 | |||
| o3 | o3-2025-04-16_low | o3 (low) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | Unknown | OpenAI's latest full-size o-series reasoning model release, evaluated on low reasoning effort. | o3 | o3 | 147.27 | 144.19 | 154.26 | |||
| o4-mini | o4-mini-2025-04-16_low | o4-mini (low) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | Unknown | The latest smaller size o-series reasoning model from OpenAI, evaluated with low reasoning effort. | o4-mini | o4-mini | 146.91 | 143.73 | 153.62 | |||
| o4-mini | o4-mini-2025-04-16_medium | o4-mini (medium) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | Unknown | The latest smaller size o-series reasoning model from OpenAI, evaluated with medium reasoning effort. | o4-mini | o4-mini | 146.91 | 143.73 | 153.62 | |||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_16K | Claude Sonnet 4.5 (16k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Anthropic's latest mid-size model, part of the Claude 4.5 series, evaluated with up to 16,000 reasoning tokens. | Claude Sonnet 4.5 (16k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 147.15 | 144.42 | 153.64 | ||
| GPT-4 (Jun 2023) | gpt-4-0613 | GPT-4 (Jun 2023) | OpenAI | United States of America | 2023-06-13 | API access | 2023-03-15 | 2.1e+25 | Likely | The June 2023 version of GPT-4, OpenAI's 2023 multimodal flagship model. | GPT-4 (Jun 2023) | gpt-4-jun-2023 | 121.80 | 119.11 | 125.67 | ||
| GPT-4 (Mar 2023) | gpt-4-0314 | GPT-4 (Mar 2023) | OpenAI | United States of America | 2023-03-14 | API access | 2023-03-15 | 2.1e+25 | Likely | The March 2023 version of GPT-4, OpenAI's 2023 multimodal flagship model. | GPT-4 (Mar 2023) | gpt-4-mar-2023 | 125.96 | 121.50 | 133.96 | ||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001 | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 | claude-haiku-4-5 | 143.09 | 140.17 | 149.33 | |||||
| Grok 4 Heavy | grok-4-heavy-web-app | Grok 4 Heavy (web app) | xAI | United States of America | 2025-07-10 | Unreleased | 2025-07-10 | Unknown | Grok 4 Heavy | grok-4-heavy | |||||||
| GPT-5 Pro | gpt-5-pro-2025-10-06_high | OpenAI | United States of America | 2025-10-07 | API access | 2025-10-07 | Unknown | GPT-5 Pro | gpt-5-pro | 150.21 | 146.92 | 157.69 | |||||
| GLM-4.5 | glm-4.5 | Z.ai (Zhipu AI),Tsinghua University | China | 2025-08-03 | Open weights (unrestricted) | 2025-08-05 | 4.4e+24 | Confident | Zhipu AI & Tsinghua University’s large open source model. | GLM-4.5 | glm-4-5 | ||||||
| GPT-5 nano | gpt-5-nano-2025-08-07_medium | GPT-5 nano (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 nano (medium) | GPT-5 nano | gpt-5-nano | 140.95 | 137.13 | 147.50 | ||
| Claude Opus 4 | claude-opus-4-20250514_27K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 27,000 reasoning tokens. | Claude Opus 4 (27k thinking) | Claude Opus 4 | claude-opus-4 | 143.37 | 140.42 | 150.37 | ||
| Claude Opus 4.1 | claude-opus-4-1-20250805_16K | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model, benchmarked with up to 16,000 reasoning tokens. | Claude Opus 4.1 (16k thinking) | Claude Opus 4.1 | claude-opus-4-1 | 144.73 | 141.67 | 150.99 | |||
| GPT-4o mini | gpt-4o-mini-2024-07-18 | OpenAI | United States of America | 2024-07-18 | API access | 2024-07-18 | Speculative | A smaller version of OpenAI's GPT-4o. | GPT-4o mini | gpt-4o-mini | 127.04 | 122.63 | 130.91 | ||||
| Claude Sonnet 4 | claude-sonnet-4-20250514 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized version of Claude 4 series of reasoning models. | Claude Sonnet 4 | Claude Sonnet 4 | claude-sonnet-4 | 142.57 | 138.67 | 149.90 | |||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro-preview-06-05 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America | 2025-06-05 | API access | 2025-06-05 | Unknown | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model. | Gemini 2.5 Pro Preview (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.72 | 143.80 | 153.41 | |
| Claude 3.5 Sonnet (October 2024) | claude-3-5-sonnet-20241022 | Claude 3.5 Sonnet (Oct 2024) | Anthropic | United States of America | 2024-10-22 | API access | 2024-10-22 | Unverified | An updated version of the top model in the Claude 3.5 series, a previous flagship model series from Anthropic. | Claude 3.5 Sonnet (Oct 2024) | Claude 3.5 Sonnet (October 2024) | claude-3-5-sonnet-october-2024 | 134.37 | 130.57 | 142.54 | ||
| Claude 3.5 Sonnet | claude-3-5-sonnet-20240620 | Claude 3.5 Sonnet (Jun 2024) | Anthropic | United States of America | 2024-06-20 | API access | 2024-06-20 | 2.7e+25 | Speculative | The first version of the top model in the Claude 3.5 series, a previous flagship model series from Anthropic. | Claude 3.5 Sonnet (Jun 2024) | Claude 3.5 Sonnet | claude-3-5-sonnet | 130.00 | 127.56 | 135.64 | |
| Claude Sonnet 4 | claude-sonnet-4-20250514_59K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 59,000 reasoning tokens. | Claude Sonnet 4 (59k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.57 | 138.67 | 149.90 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_64K | Claude 3.7 Sonnet (64k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 64,000 reasoning tokens allowed. | Claude 3.7 Sonnet (64k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 142.03 | 139.52 | 148.45 | |
| Grok 3 | grok-3-beta | xAI | United States of America | 2025-04-09 | API access | 2025-02-17 | 3.5e+26 | Likely | A beta release of XAI's third generation Grok model. | Grok 3 (beta) | Grok 3 | grok-3 | 139.18 | 136.31 | 144.98 | ||
| Magistral Small 1.1 | magistral-small-2506 | Mistral AI | France | 2025-06-10 | Open weights (unrestricted) | 2025-06-10 | Confident | Magistral Small 1.1 | magistral-small-1-1 | 133.18 | 128.92 | 137.84 | |||||
| Qwen3-235B-A22B | qwen3-235b-a22b | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Likely | Qwen3 (Apr 2025) | Qwen3-235B-A22B | qwen3-235b-a22b | 139.73 | 135.99 | 145.21 | |||
| Gemini 2.5 Pro (May 2025) | gemini-2.5-pro-preview-05-06 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America | 2025-05-06 | API access | 2025-05-06 | Unknown | A May 2025 preview version of Gemini 2.5 Pro, a flagship reasoning model from Google DeepMind. | Gemini 2.5 Pro Preview (Jun 2025) | Gemini 2.5 Pro (May 2025) | Gemini 2.5 Pro (May 2025) | gemini-2-5-pro-may-2025 | 142.80 | 139.26 | 149.39 | |
| DeepSeek-R1 | DeepSeek-R1 | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-20 | 3.5e+24 | Confident | DeepSeek-R1 | DeepSeek-R1 | deepseek-r1 | 139.86 | 137.43 | 145.18 | |||
| Claude Sonnet 4 | claude-sonnet-4-20250514_32K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 32,000 reasoning tokens. | Claude Sonnet 4 (32k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.57 | 138.67 | 149.90 | |||
| Claude Opus 4 | claude-opus-4-20250514_16K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 16,000 reasoning tokens. | Claude Opus 4 (16k thinking) | Claude Opus 4 | claude-opus-4 | 143.37 | 140.42 | 150.37 | ||
| Claude Sonnet 4 | claude-sonnet-4-20250514_16K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 16,000 reasoning tokens. | Claude Sonnet 4 (16k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.57 | 138.67 | 149.90 | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20 | Google DeepMind | United States of America | 2025-05-20 | API access | 2025-04-17 | Unknown | The May 2025 preview version of Gemini 2.5 Flash, a small reasoning model from Google DeepMind's Gemini 2.5 series. | Gemini 2.5 Flash (May 2025) | Gemini 2.5 Flash (May 2025) | gemini-2-5-flash-may-2025 | 142.45 | 140.05 | 148.50 | |||
| Qwen Plus | qwen-plus-2025-04-28 | Alibaba | China | 2025-04-28 | API access | 2024-02-06 | Unknown | Qwen Plus (Apr 2025) | Qwen Plus (Apr 2025) | Qwen Plus (Apr 2025) | qwen-plus-apr-2025 | ||||||
| DeepSeek-V3 (Mar 2025) | DeepSeek-V3-0324 | DeepSeek-V3 (Mar 2025) | DeepSeek | China | 2025-03-24 | Open weights (restricted use) | 2025-03-24 | 3.3e+24 | Confident | An updated version of the DeepSeek V3 model. | DeepSeek-V3 (Mar 2025) | DeepSeek-V3 (Mar 2025) | deepseek-v3-mar-2025 | 137.19 | 134.71 | 142.64 | |
| Gemini 2.5 Pro (Mar 2025) | gemini-2.5-pro-preview-03-25 | Gemini 2.5 Pro Preview (Mar 2025) | Google DeepMind | United States of America | 2025-03-31 | API access | 2025-03-25 | Unknown | A March 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship model. | Gemini 2.5 Pro Preview (Mar 2025) | Gemini 2.5 Pro (Mar 2025) | Gemini 2.5 Pro (Mar 2025) | gemini-2-5-pro-mar-2025 | 145.07 | 142.02 | 152.29 | |
| Mistral Medium 3 | mistral-medium-2505 | Mistral AI | France | 2025-05-07 | API access | 2025-05-07 | Unknown | A 2025 language model from Mistral. | Mistral Medium 3 | mistral-medium-3 | 135.43 | 126.46 | 140.41 | ||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-04-17 | Gemini 2.5 Flash Preview (Apr 2025) | Google DeepMind | United States of America | 2025-04-17 | API access | 2025-04-17 | Unknown | An April 2025 preview of Gemini 2.5 Flash, a small model in Google DeepMind's Gemini 2.5 series. | Gemini 2.5 Flash (Apr 2025) | Gemini 2.5 Flash (Apr 2025) | gemini-2-5-flash-apr-2025 | 140.71 | 137.20 | 146.83 | ||
| GPT-4.1 nano | gpt-4.1-nano-2025-04-14 | GPT-4.1 nano | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | Unknown | The smallest version of OpenAI's GPT-4.1. | GPT-4.1 nano | gpt-4-1-nano | 130.92 | 127.67 | 136.23 | |||
| GPT-4.1 mini | gpt-4.1-mini-2025-04-14 | GPT-4.1 mini | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | Unknown | A smaller version of OpenAI's GPT-4.1. | GPT-4.1 mini | gpt-4-1-mini | 135.87 | 132.68 | 140.65 | |||
| QWQ-Plus | qwq-plus | Alibaba | China | 2025-04-08 | API access | 2025-04-08 | Unknown | QwQ-Plus | QWQ-Plus | qwq-plus | |||||||
| Grok-3 mini | grok-3-mini-beta_low | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | Unknown | A beta release of a smaller version of XAI's third generation Grok model, evaluated with low reasoning effort. | Grok-3 mini | grok-3-mini | 141.19 | 138.81 | 147.60 | ||||
| Llama 4 Maverick | Llama-4-Maverick-17B-128E-Instruct-FP8 | Llama 4 Maverick (FP8) | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | Likely | The 17 billion active (400 billion total)-parameter model in Meta's Llama 4 series, quantized to FP8. | Llama 4 Maverick | llama-4-maverick | 133.13 | 129.73 | 138.87 | ||
| Llama 4 Scout | Llama-4-Scout-17B-16E-Instruct | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 4.1e+24 | Likely | The 17 billion active (109 billion total)-parameter model in Meta's Llama 4 series. | Llama 4 Scout | llama-4-scout | 130.61 | 127.27 | 135.54 | |||
| Qwen-Turbo | qwen-turbo-2024-11-01 | Alibaba | China | 2024-11-01 | API access | 2024-02-06 | Unknown | Qwen Turbo | Qwen-Turbo | qwen-turbo | |||||||
| Qwen Plus | qwen-plus-2025-01-25 | Alibaba | China | 2025-01-25 | API access | 2024-02-06 | Unknown | Qwen Plus (Jan 2025) | Qwen Plus (Jan 2025) | Qwen Plus (Jan 2025) | qwen-plus-jan-2025 | ||||||
| Qwen2.5-Max | qwen-max-2025-01-25 | Alibaba | China | 2025-01-25 | API access | 2025-01-28 | Unknown | Qwen2.5-Max | Qwen2.5-Max | qwen2-5-max | 133.49 | 129.89 | 139.94 | ||||
| Hermes 2 Theta Llama-3 70B | Hermes-2-Theta-Llama-3-70B | Nous Research,Arcee AI | United States of America | 2024-06-20 | Open weights (restricted use) | 2024-06-20 | Confident | Nous Research’s fine-tuned Llama-3 70B optimized for instruction following and chat. | Hermes 2 Theta Llama-3 70B | hermes-2-theta-llama-3-70b | |||||||
| Gemini 2.5 Pro (Mar 2025) | gemini-2.5-pro-exp-03-25 | Gemini 2.5 Pro Exp (Mar 2025) | Google DeepMind | United States of America | 2025-03-25 | API access | 2025-03-25 | Unknown | A March 2025 preview version of Gemini 2.5 Pro. | Gemini 2.5 Pro Exp (Mar 2025) | Gemini 2.5 Pro (Mar 2025) | Gemini 2.5 Pro (Mar 2025) | gemini-2-5-pro-mar-2025 | 145.07 | 142.02 | 152.29 | |
| Mistral Small 3 | mistral-small-2501 | Mistral AI | France | 2025-01-25 | Open weights (unrestricted) | 2025-01-30 | 1.2e+24 | Confident | Mistral Small 3 | mistral-small-3 | |||||||
| Mistral Small 3.1 | mistral-small-2503 | Mistral AI | France | 2025-03-17 | Open weights (unrestricted) | 2025-03-17 | Confident | An updated version of Mistral small, a 24 billion-parameter model from Mistral. | Mistral Small 3.1 | mistral-small-3-1 | |||||||
| Gemma 3 27B | gemma-3-27b-it | Google DeepMind | United States of America | 2025-03-12 | Open weights (restricted use) | 2025-03-12 | 2.3e+24 | Confident | An instruction-tuned version of the 27 billion-parameter model Google DeepMind's Gemma 3 series. | Gemma 3 27B | gemma-3-27b | 131.14 | 127.45 | 136.64 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_32K | Claude 3.7 Sonnet (32k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 32,000 reasoning tokens allowed. | Claude 3.7 Sonnet (32k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 142.03 | 139.52 | 148.45 | |
| Gemini 1.5 Flash 8B | gemini-1.5-flash-8b-001 | Google DeepMind | United States of America | 2024-10-03 | API access | 2024-05-10 | Confident | Gemini 1.5 Flash 8B | gemini-1-5-flash-8b | ||||||||
| DeepSeek-R1-Distill-Qwen-14B | DeepSeek-R1-Distill-Qwen-14B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | Speculative | A model based on Qwen 14B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | DeepSeek-R1-Distill-Qwen-14B | deepseek-r1-distill-qwen-14b | |||||||
| DeepSeek-R1-Distill-Llama-70B | DeepSeek-R1-Distill-Llama-70B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | Speculative | A model based on LLaMA 70B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | DeepSeek-R1-Distill-Llama-70B | deepseek-r1-distill-llama-70b | |||||||
| Gemini 2.0 Flash | gemini-2.0-flash-001 | Gemini 2.0 Flash (Feb 2025) | Google DeepMind,Google | United States of America | 2025-02-05 | API access | 2024-12-11 | Unknown | The first version of Google DeepMind's Gemini 2.0 Flash model. | Gemini 2.0 Flash (Feb 2025) | Gemini 2.0 Flash (Feb 2025) | gemini-2-0-flash-feb-2025 | 135.89 | 133.78 | 142.01 | ||
| DeepSeek-V3 | DeepSeek-V3 | DeepSeek | China | 2024-12-26 | Open weights (restricted use) | 2024-12-24 | 3.3e+24 | Confident | DeepSeek’s 2024 mixture-of-experts model. | DeepSeek-V3 | DeepSeek-V3 | deepseek-v3 | 133.13 | 130.21 | 137.90 | ||
| Tulu 3 (Tülu 3) 70B | Llama-3.1-Tulu-3-70B-DPO | Allen Institute for AI,University of Washington | United States of America | 2024-11-21 | Open weights (restricted use) | 2024-11-21 | Confident | Tülu 3 70B | Tulu 3 (Tülu 3) 70B | tulu-3-tlu-3-70b | |||||||
| Gemma 2 27B | gemma-2-27b-it | Google DeepMind | United States of America | 2024-06-24 | Open weights (restricted use) | 2024-06-24 | 2.1e+24 | Confident | An instruction-optimized version of Google DeepMind's Gemma 2 27B. | Gemma 2 27B | gemma-2-27b | 122.48 | 119.24 | 126.44 | |||
| Gemma 2 9B | gemma-2-9b-it | Google DeepMind | United States of America | 2024-06-24 | Open weights (restricted use) | 2024-06-24 | 4.3e+23 | Confident | Gemma 2 9B | gemma-2-9b | 119.51 | 116.30 | 122.43 | ||||
| Claude 2.1 | claude-2.1 | Anthropic | United States of America | 2023-11-21 | API access | 2023-11-21 | Unknown | An updated version in Anthropic's Claude 2 series. | Claude 2.1 | claude-2-1 | 118.20 | 114.11 | 123.14 | ||||
| Claude 2 | claude-2.0 | Anthropic | United States of America | 2023-07-11 | API access | 2023-07-11 | 3.9e+24 | Speculative | Anthropic's second generation Claude model. | Claude 2 | claude-2 | 119.53 | 115.97 | 126.66 | |||
| Gemini 2.0 Flash Thinking | gemini-2.0-flash-thinking-exp-01-21 | Gemini 2.0 Flash Thinking Exp | Google DeepMind,Google | United States of America | 2025-01-21 | API access | 2024-12-19 | Unknown | A January 2025 experimental version of Gemini 2.0 Flash Thinking, a small reasoning model from Google DeepMind. | Gemini 2.0 Flash Thinking (Jan 2025) | Gemini 2.0 Flash Thinking (Jan 2025) | gemini-2-0-flash-thinking-jan-2025 | 136.39 | 131.85 | 142.24 | ||
| o1-preview | o1-preview-2024-09-12 | o1-preview | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | Unknown | The September 2024 preview version of OpenAI’s first reasoning model, o1. | o1-preview | o1-preview | 135.97 | 132.94 | 143.08 | |||
| Gemini 2.0 Pro | gemini-2.0-pro-exp-02-05 | Gemini 2.0 Pro Exp (Feb 2025) | Google DeepMind | United States of America | 2025-02-05 | Hosted access (no API) | 2024-12-11 | Unknown | A February 2025 experimental version of Google DeepMind's previous flagship model, Gemini 2.0 Pro. | Gemini 2.0 Pro | gemini-2-0-pro | 135.75 | 132.29 | 142.18 | |||
| o1 | o1-2024-12-17_high | o1 (high) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | Unknown | OpenAI's first reasoning model, evaluated with at the high reasoning effort level. | o1 | o1 | 142.83 | 139.83 | 148.65 | |||
| GPT-4o | gpt-4o-2024-08-06 | GPT-4o (Aug 2024) | OpenAI | United States of America | 2024-08-06 | API access | 2024-05-13 | Speculative | An updated version of GPT-4o, OpenAI's model that powered ChatGPT from mid-2024 to -2025. | GPT-4o (Aug 2024) | GPT-4o (Aug 2024) | gpt-4o-aug-2024 | 129.27 | 126.95 | 133.46 | ||
| Gemini 1.5 Flash | gemini-1.5-flash-002 | Google DeepMind | United States of America | 2024-09-24 | API access | 2024-05-10 | Unknown | The second version of Google DeepMind's Gemini 1.5 Flash. | Gemini 1.5 Flash (Sep 2024) | Gemini 1.5 Flash (Sep 2024) | gemini-1-5-flash-sep-2024 | 130.44 | 125.87 | 134.91 | |||
| o1-mini | o1-mini-2024-09-12_high | o1-mini (high) | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | Unknown | A smaller version of OpenAI’s first reasoning model, o1, evaluated with high reasoning effort. | o1-mini | o1-mini | 136.68 | 133.18 | 142.63 | |||
| o1-mini | o1-mini-2024-09-12_medium | o1-mini (medium) | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | Unknown | A smaller version of OpenAI’s first reasoning model, o1, evaluated with medium reasoning effort. | o1-mini | o1-mini | 136.68 | 133.18 | 142.63 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_16K | Claude 3.7 Sonnet (16k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 16,000 reasoning tokens allowed. | Claude 3.7 Sonnet (16k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 142.03 | 139.52 | 148.45 | |
| o3-mini | o3-mini-2025-01-31_medium | o3-mini (medium) | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | Unknown | A smaller version of OpenAI's o3 reasoning model, evaluated with medium reasoning effort. | o3-mini | o3-mini | 141.63 | 138.79 | 147.93 | |||
| Grok-2 | grok-2-1212 | xAI | United States of America | 2024-12-12 | API access | 2024-08-13 | 3.0e+25 | Confident | XAI's second generation Grok model. | Grok-2 (Dec 2024) | Grok-2 (Dec 2024) | grok-2-dec-2024 | 130.84 | 128.22 | 135.32 | ||
| Mistral Large 2 | mistral-large-2411 | Mistral AI | France | 2024-11-18 | Open weights (non-commercial) | 2024-07-24 | 2.1e+25 | Likely | Mistral Large 2 (Nov 2024) | Mistral Large 2 (Nov 2024) | mistral-large-2-nov-2024 | 128.85 | 125.43 | 133.47 | |||
| GPT-4.5 | gpt-4.5-preview-2025-02-27 | GPT-4.5 Preview (Feb 2025) | OpenAI | United States of America | 2025-02-27 | API access | 2025-02-27 | 3.8e+26 | Likely | The largest model in OpenAI’s GPT series. | GPT-4.5 | gpt-4-5 | 137.71 | 135.12 | 142.72 | ||
| GPT-4 Turbo (Apr 2024) | gpt-4-turbo-2024-04-09 | OpenAI | United States of America | 2024-04-09 | API access | 2024-04-09 | Unknown | The April 2024 version of GPT-4 Turbo, OpenAI's then-flagship language model. | GPT-4 Turbo (Apr 2024) | gpt-4-turbo-apr-2024 | 127.64 | 125.16 | 132.30 | ||||
| o1 | o1-2024-12-17_medium | o1 (medium) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | Unknown | OpenAI's first reasoning model, evaluated with at the medium reasoning effort level. | o1 | o1 | 142.83 | 139.83 | 148.65 | |||
| Llama 3-70B | Meta-Llama-3-70B-Instruct | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.9e+24 | Confident | An instruction-tuned version of the 70 billion-parameter model in Meta’s LLaMA 3 series. | Llama 3-70B | llama-3-70b | 122.65 | 120.47 | 126.55 | |||
| Gemini 1.5 Pro | gemini-1.5-pro-001 | Google DeepMind | United States of America | 2024-05-14 | API access | 2024-02-15 | Speculative | The first of two versions of a previous flagship model from Google DeepMind, Gemini 1.5 Pro. | Gemini 1.5 Pro (May 2024) | Gemini 1.5 Pro (May 2024) | gemini-1-5-pro-may-2024 | 127.28 | 125.04 | 133.57 | |||
| Llama 3.3 70B | Llama-3.3-70B-Instruct | Meta AI | United States of America | 2024-12-06 | Open weights (restricted use) | 2024-12-06 | 6.9e+24 | Confident | An instruction-tuned version of the 70 billion-parameter model in Meta's Llama 3.3 series. | Llama 3.3 70B | llama-3-3-70b | 127.53 | 124.74 | 132.61 | |||
| Llama 3.2 90B | Llama-3.2-90B-Vision-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | Confident | An instruction-tuned version of the 90 billion-parameter vision model in Meta's Llama 3.3 series. | Llama 3.2 90B | llama-3-2-90b | 125.74 | 122.94 | 129.96 | ||||
| Llama 2-70B | Llama-2-70b-chat-hf | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | Confident | A chat-optimized version of the 70-billion parameter model in Meta's Llama 2 series. | Llama 2-70B | llama-2-70b | 113.14 | 109.14 | 116.76 | |||
| Gemini 1.5 Pro | gemini-1.5-pro-002 | Google DeepMind | United States of America | 2024-09-24 | API access | 2024-02-15 | Speculative | The first of two versions of a previous flagship model from Google DeepMind, Gemini 1.5 Pro. | Gemini 1.5 Pro (Sept 2024) | Gemini 1.5 Pro (Sept 2024) | gemini-1-5-pro-sept-2024 | 132.83 | 130.38 | 137.88 | |||
| Qwen2.5-32B | qwen2.5-32b-instruct | Alibaba | China | 2024-09-17 | Open weights (unrestricted) | 2024-09-17 | 3.5e+24 | Confident | Qwen2.5-32B (instruct) | Qwen2.5-32B | qwen2-5-32b | ||||||
| Mistral Large | mistral-large-2402 | Mistral AI | France | 2024-02-26 | API access | 2024-02-26 | 1.1e+25 | Likely | Mistral Large | mistral-large | 121.05 | 115.53 | 126.48 | ||||
| Claude 3 Sonnet | claude-3-sonnet-20240229 | Anthropic | United States of America | 2024-02-29 | API access | 2024-03-04 | Unknown | The mid-sized model in Anthropic's Claude 3 model series. | Claude 3 Sonnet | Claude 3 Sonnet | claude-3-sonnet | 119.89 | 112.29 | 124.46 | |||
| Gemini 1.0 Pro | gemini-1.0-pro-001 | Google DeepMind | United States of America | 2023-12-13 | API access | 2023-12-06 | Speculative | An earlier flagship model from Google DeepMind. | Gemini 1.0 Pro | gemini-1-0-pro | 116.93 | 114.86 | 121.56 | ||||
| Llama 3.1-8B | Llama-3.1-8B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 1.2e+24 | Likely | An instruction-tuned version of the 8 billion-parameter model in Meta’s LLaMA 3.1 series. | Llama 3.1-8B | llama-3-1-8b | 115.51 | 104.55 | 120.99 | |||
| Llama 3.1-405B | Llama-3.1-405B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 3.8e+25 | Confident | An instruction-tuned version of the 405 billion-parameter model in Meta’s LLaMA 3.1 series. | Llama 3.1-405B | llama-3-1-405b | 129.05 | 126.47 | 133.63 | |||
| Qwen2.5-72B | qwen2.5-72b-instruct | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | Confident | An instruction-tuned version of the 72 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-72B | Qwen2.5-72B | qwen2-5-72b | 129.40 | 124.88 | 133.16 | ||
| GPT-4o | gpt-4o-2024-05-13 | GPT-4o (May 2024) | OpenAI | United States of America | 2024-05-13 | API access | 2024-05-13 | Speculative | The first version of GPT-4o, OpenAI's last multimodal model in the GPT-4 series, which powered ChatGPT from mid-2024 to -2025. | GPT-4o (May 2024) | GPT-4o (May 2024) | gpt-4o-may-2024 | 128.91 | 126.30 | 134.14 | ||
| Llama 3.1-70B | Llama-3.1-70B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 7.9e+24 | Confident | An instruction-optimized version of Llama 3.1 70B, the 70 billion-parameter model in Meta's Llama 3.1 series. | Llama 3.1-70B | llama-3-1-70b | 125.52 | 122.50 | 130.63 | |||
| Gemini 1.5 Flash | gemini-1.5-flash-001 | Gemini 1.5 Flash (May 2024) | Google DeepMind | United States of America | 2024-05-23 | API access | 2024-05-10 | Unknown | The first version of Google DeepMind's Gemini 1.5 Flash. | Gemini 1.5 Flash (May 2024) | Gemini 1.5 Flash (May 2024) | gemini-1-5-flash-may-2024 | 122.52 | 118.82 | 126.23 | ||
| Claude 3 Haiku | claude-3-haiku-20240307 | Anthropic | United States of America | 2024-03-07 | API access | 2024-03-04 | Unknown | The smallest model in Anthropic's Claude 3 series. | Claude 3 Haiku | Claude 3 Haiku | claude-3-haiku | 117.57 | 113.07 | 121.00 | |||
| Llama 3-8B | Meta-Llama-3-8B-Instruct | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.2e+23 | Confident | An instruction-tuned version of the 8 billion-parameter model in Meta's Llama 3 series. | Llama 3-8B | llama-3-8b | 115.92 | 112.53 | 118.72 | |||
| Claude 3 Opus | claude-3-opus-20240229 | Anthropic | United States of America | 2024-02-29 | API access | 2024-03-04 | Speculative | The largest model in Anthropic's Claude 3 model series. | Claude 3 Opus | Claude 3 Opus | claude-3-opus | 126.97 | 124.33 | 131.88 | |||
| Phi-4 | phi-4 | Microsoft Research | United States of America | 2024-12-12 | Open weights (unrestricted) | 2024-12-12 | 9.3e+23 | Confident | Phi-4 | Phi-4 | phi-4 | 131.15 | 128.00 | 135.78 | |||
| Mistral Large 2 | mistral-large-2407 | Mistral AI | France | 2024-07-24 | Open weights (non-commercial) | 2024-07-24 | 2.1e+25 | Likely | Mistral Large 2 (Jul 2024) | Mistral Large 2 (Jul 2024) | mistral-large-2-jul-2024 | 127.60 | 125.34 | 132.47 | |||
| phi-3-medium 14B | Phi-3-medium-128k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 4.0e+23 | Likely | phi-3-medium 14B | phi-3-medium-14b | 120.82 | 117.99 | 125.14 | ||||
| GPT-4 Turbo (Nov 2023) | gpt-4-0125-preview | GPT-4 Turbo Preview (January 2024) | OpenAI | United States of America | 2024-01-25 | API access | 2023-11-06 | Unknown | A January 2024 preview version of GPT-4 Turbo, OpenAI's updated GPT-4 series model. | GPT-4 Turbo (Nov 2023) | gpt-4-turbo-nov-2023 | ||||||
| GPT-3.5 Turbo | gpt-3.5-turbo-1106 | GPT-3.5 Turbo (Nov 2023) | OpenAI | United States of America | 2023-11-06 | API access | 2023-06-13 | Likely | A November 2023 preview of GPT-3.5 turbo, OpenAI's updated GPT-3.5 series model. | GPT-3.5 Turbo (Nov 2023) | GPT-3.5 Turbo (Nov 2023) | gpt-3-5-turbo-nov-2023 | 117.87 | 110.16 | 122.42 | ||
| GPT-4 Turbo (Nov 2023) | gpt-4-1106-preview | GPT-4 Turbo Preview (Nov 2023) | OpenAI | United States of America | 2023-11-06 | API access | 2023-11-06 | Unknown | A November 2023 preview of GPT-4 turbo, OpenAI's updated GPT-4 series model. | GPT-4 Turbo (Nov 2023) | gpt-4-turbo-nov-2023 | ||||||
| GPT-3.5 Turbo | gpt-3.5-turbo-0125 | GPT-3.5 Turbo (Jan 2024) | OpenAI | United States of America | 2024-01-25 | API access | 2023-06-13 | Likely | A January 2024 preview version of GPT-3.5 Turbo, OpenAI's updated GPT-3.5 series model. | GPT-3.5 Turbo (Jan 2024) | GPT-3.5 Turbo (Jan 2024) | gpt-3-5-turbo-jan-2024 | 114.16 | 106.02 | 118.33 | ||
| Yi-1.5-34B | Yi-1.5-34B-Chat | 01.AI | China | 2024-05-13 | Open weights (restricted use) | 2024-05-13 | 7.3e+23 | Confident | Yi-1.5-34B (chat) | Yi-1.5-34B | yi-1-5-34b | ||||||
| Yi-34B | Yi-34B-Chat | 01.AI | China | 2023-11-22 | Open weights (restricted use) | 2023-11-02 | 6.1e+23 | Confident | Yi-34B (chat) | Yi-34B | Yi-34B | yi-34b | 117.40 | 114.48 | 121.38 | ||
| Qwen2-72B | qwen2-72b-instruct | Alibaba | China | 2024-06-07 | Open weights (unrestricted) | 2024-06-07 | 3.0e+24 | Confident | Qwen's 72 billion-parameter model in the Qwen 2 series. | Qwen2-72B | Qwen2-72B | qwen2-72b | 125.71 | 122.38 | 129.93 | ||
| Qwen1.5-72B | qwen1.5-72b-chat | Alibaba | China | 2024-02-04 | Open weights (restricted use) | 2024-02-04 | 1.3e+24 | Confident | Qwen1.5-72B | Qwen1.5-72B | qwen1-5-72b | ||||||
| Qwen1.5-32B | qwen1.5-32b-chat | Alibaba | China | 2024-04-03 | Open weights (restricted use) | 2024-02-05 | Confident | A chat-optimized 32 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-32B (chat) | Qwen1.5-32B | qwen1-5-32b | ||||||
| Mistral 7B | Mistral-7B-Instruct-v0.3 | Mistral AI | France | 2024-05-27 | Open weights (unrestricted) | 2023-10-10 | Confident | An instruction-tuned version of Mistral-7B-v0.3, an updated version of their 7 billion-parameter model. | Mistral 7B v0.3 | Mistral 7B v0.3 | mistral-7b-v0-3 | ||||||
| DeepSeek LLM 67B | deepseek-llm-67b-chat | DeepSeek | China | 2023-11-29 | Open weights (restricted use) | 2024-01-05 | 8.0e+23 | Confident | DeepSeek LLM 67B (chat) | DeepSeek LLM 67B | deepseek-llm-67b | ||||||
| Mixtral 8x7B | Mixtral-8x7B-Instruct-v0.1 | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | Speculative | An instruction-tuned version of Mistral's 7 billion active (56 billion total)-parameter mixture-of-experts model | Mixtral 8x7B | mixtral-8x7b | 117.71 | 114.62 | 121.42 | |||
| WizardLM-2 8x22B | WizardLM-2-8x22B | Microsoft | United States of America | 2024-04-15 | Open weights (unrestricted) | 2024-04-15 | Confident | WizardLM-2 8x22B | WizardLM-2 8x22B | wizardlm-2-8x22b | |||||||
| DBRX | dbrx-instruct | Databricks | United States of America | 2024-03-27 | Open weights (restricted use) | 2024-03-27 | 2.6e+24 | Confident | DBRX (instruct) | DBRX | dbrx | ||||||
| Ministral 3B | ministral-3b-2410 | Mistral AI | France | 2024-10-16 | API access | 2024-10-16 | Confident | Ministral 3B | ministral-3b | ||||||||
| Mixtral 8x22B | open-mixtral-8x22b | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | Speculative | Mixtral 8x22B | mixtral-8x22b | 121.21 | 114.22 | 125.22 | ||||
| Ministral 8B | ministral-8b-2410 | Mistral AI | France | 2024-10-16 | Open weights (non-commercial) | 2024-10-16 | Confident | Ministral 8B | ministral-8b | ||||||||
| Mixtral 8x7B | open-mixtral-8x7b | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | Speculative | Mixtral 8x7B | mixtral-8x7b | 117.71 | 114.62 | 121.42 | ||||
| Mistral 7B | open-mistral-7b | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | Confident | Mistral 7B v0.3 | Mistral 7B v0.3 | mistral-7b-v0-3 | |||||||
| Mistral NeMo | open-mistral-nemo-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | Confident | Mistral NeMo | mistral-nemo | 118.27 | 114.37 | 123.07 | |||||
| Eurus-2-7B-PRIME | Eurus-2-7B-PRIME | Tsinghua University,University of Illinois Urbana-Champaign (UIUC),Shanghai AI Lab,Peking University,Shanghai Jiao Tong University,CUHK Shenzhen Research Institute | United States of America,China | 2024-12-31 | Open weights (unrestricted) | 2025-02-03 | Speculative | Eurus-2-7B-PRIME | eurus-2-7b-prime | ||||||||
| Gemini 2.5 Deep Think | gemini-2.5-deep-think-2025-08-01-webapp | Google,Google DeepMind | United States of America | 2025-08-01 | Hosted access (no API) | 2025-08-01 | Unknown | Gemini 2.5 Deep Think | gemini-2-5-deep-think | ||||||||
| xiaoyi-deepresearch | 2026-03-06 | Unverified | xiaoyi-deepresearch | ||||||||||||||
| GPT-5 | gpt-5-2025-08-07_unknown | GPT-5 (unknown thinking) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 (high) | GPT-5 | gpt-5 | 150.00 | 146.38 | 157.62 | |
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_2K | Claude Sonnet 4.5 (2k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (2k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 147.15 | 144.42 | 153.64 | |||
| GPT-5 | gpt-5-2025-08-07_low | GPT-5 (low) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using low reasoning effort. | GPT-5 (low) | GPT-5 | gpt-5 | 150.00 | 146.38 | 157.62 | |
| GPT-5 | gpt-5-2025-08-07_minimal | GPT-5 (minimal) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using minimal reasoning effort. | GPT-5 (minimal) | GPT-5 | gpt-5 | 150.00 | 146.38 | 157.62 | |
| Claude Opus 4 | claude-opus-4-20250514_2K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 2,000 reasoning tokens. | Claude Opus 4 (2k thinking) | Claude Opus 4 | claude-opus-4 | 143.37 | 140.42 | 150.37 | ||
| Claude Sonnet 4 | claude-sonnet-4-20250514_2K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized Claude 4 series model, evaluated at up to 2,000 reasoning tokens. | Claude Sonnet 4 (2k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.57 | 138.67 | 149.90 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_2K | Claude 3.7 Sonnet (2k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 2,000 reasoning tokens allowed. | Claude 3.7 Sonnet (2k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 142.03 | 139.52 | 148.45 | |
| Sonar | sonar-pro | Perplexity | United States of America | 2025-01-21 | API access | 2025-02-11 | Confident | A search-optimized language model from Perplexity. | Sonar Pro | Sonar | sonar | ||||||
| sonar | Perplexity Sonar | Unverified | sonar | ||||||||||||||
| Amazon Nova Pro | amazon.nova-pro-v1:0 | Amazon | United States of America | 2024-12-03 | API access | 2024-12-03 | 6.0e+24 | Speculative | Amazon Nova Pro | amazon-nova-pro | |||||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-09-2025 | Google DeepMind | United States of America | 2025-09-25 | API access | 2025-04-17 | Unknown | Gemini 2.5 Flash (Sep 2025) | Gemini 2.5 Flash (Sep 2025) | Gemini 2.5 Flash (Sep 2025) | gemini-2-5-flash-sep-2025 | 143.39 | 139.01 | 151.42 | |||
| DeepSeek-V3.2-Exp | DeepSeek-V3.2-Exp | DeepSeek | China | 2025-09-29 | Open weights (unrestricted) | 2025-09-29 | 4.2e+24 | Confident | DeepSeek-V3.2-Exp | deepseek-v3-2-exp | 145.19 | 141.83 | 151.64 | ||||
| Claude Instant | claude-instant-1.1 | Anthropic | United States of America | API access | 2023-08-09 | Unknown | Claude Instant | claude-instant | 120.53 | 117.29 | 125.18 | ||||||
| Claude 1.3 | claude-1.3 | Anthropic | United States of America | 2023-04-18 | API access | 2023-04-18 | Unknown | An updated version of Anthropic's first generation Claude model. | Claude 1.3 | claude-1-3 | |||||||
| PaLM 2-L | 2023-05-17 | Unverified | palm-2-l | ||||||||||||||
| Llama 2-70B | Llama-2-70b-hf | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | Confident | A chat-optimized of a 70 billion-parameter model in the Llama 2 series. | Llama 2-70B | llama-2-70b | 113.14 | 109.14 | 116.76 | |||
| PaLM 2-M | 2023-05-17 | Unverified | palm-2-m | ||||||||||||||
| PaLM (540B) | PaLM 540B | Google Research | United States of America | 2022-04-04 | Unreleased | 2022-04-04 | 2.5e+24 | Confident | Google's largest model in the PaLM series. | PaLM (540B) | palm-540b | ||||||
| PaLM 2-S | 2023-05-17 | Unverified | palm-2-s | ||||||||||||||
| InstructGPT 175B | text-davinci-001 | OpenAI | United States of America | 2022-01-27 | API access | 2022-01-27 | 3.2e+23 | Confident | InstructGPT 175B | instructgpt-175b | |||||||
| Gemma 2B | gemma-2b | Google DeepMind | United States of America | 2024-02-21 | Open weights (restricted use) | 2024-02-21 | 4.5e+22 | Confident | Google's 2 billion-parameter model in the Gemma series of open source models. | Gemma 2B | gemma-2b | 93.53 | 88.58 | 97.02 | |||
| Gemma 7B | gemma-7b | Google DeepMind | United States of America | 2024-02-21 | Open weights (restricted use) | 2024-02-21 | 3.1e+23 | Confident | Google's 7 billion-parameter model in the Gemma series of open source models. | Gemma 7B | gemma-7b | 111.08 | 107.88 | 115.61 | |||
| Mistral 7B | Mistral-7B-v0.1 | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | Confident | Mistral 7B v0.1 | Mistral 7B v0.1 | mistral-7b-v0-1 | 111.36 | 108.25 | 114.82 | ||||
| Llama 2-7B | Llama-2-7b | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.4e+22 | Confident | Meta's 7 billion-parameter model in the Llama 2 series of open source models. | Llama 2-7B | llama-2-7b | 98.03 | 93.65 | 102.60 | |||
| Llama 2-13B | Llama-2-13b | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 1.6e+23 | Confident | Meta's 13 billion-parameter model in the Llama 2 series of open source models. | Llama 2-13B | llama-2-13b | 105.26 | 100.89 | 108.96 | |||
| DeepSeek-V2 (MoE-236B) | DeepSeek-V2 | DeepSeek | China | 2024-05-07 | Open weights (restricted use) | 2024-05-07 | 1.0e+24 | Confident | DeepSeek-V2 (MoE-236B, May 2024) | DeepSeek-V2 (MoE-236B, May 2024) | deepseek-v2-moe-236b-may-2024 | 124.41 | 120.71 | 128.71 | |||
| Qwen2.5-72B | Qwen2.5-72B | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | Confident | A 72 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-72B | Qwen2.5-72B | qwen2-5-72b | 129.40 | 124.88 | 133.16 | ||
| Llama 3.1-405B | Llama-3.1-405B | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 3.8e+25 | Confident | The 405 billion-parameter model in Meta’s LLaMA 3.1 series. | Llama 3.1-405B | llama-3-1-405b | 129.05 | 126.47 | 133.63 | |||
| phi-3-mini 3.8B | Phi-3-mini-4k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 7.5e+22 | Confident | A 4 billion-parameter model in Microsoft's Phi-3 open source model series with a 4000-token context length. | phi-3-mini 3.8B | phi-3-mini-3-8b | 116.62 | 111.72 | 120.24 | |||
| phi-3-small 7.4B | Phi-3-small-8k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 2.1e+23 | Confident | A 7 billion-parameter model in Microsoft's Phi-3 open source model series with a 8000-token context length. | phi-3-small 7.4B | phi-3-small-7-4b | 121.33 | 118.17 | 124.85 | |||
| Phi-2 | phi-2 | Microsoft | United States of America | 2023-12-12 | Open weights (unrestricted) | 2023-12-12 | 2.3e+22 | Confident | Phi-2 | phi-2 | 107.21 | 77.57 | 111.99 | ||||
| Mixtral 8x7B | Mixtral-8x7B-v0.1 | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | Speculative | Mixtral 8x7B | mixtral-8x7b | 117.71 | 114.62 | 121.42 | ||||
| GLaM | GLaM (MoE) | United States of America | 2021-12-13 | Unreleased | 2021-12-13 | 3.6e+23 | Confident | GLaM | glam | ||||||||
| Gopher (280B) | Gopher (280B) | DeepMind | United Kingdom of Great Britain and Northern Ireland | 2021-12-08 | Unreleased | 2021-12-08 | 6.3e+23 | Confident | A 280 billion-parameter model from DeepMind in 2021. | Gopher (280B) | gopher-280b | ||||||
| MPT-7B | mpt-7b | MosaicML | United States of America | 2023-05-05 | Open weights (unrestricted) | 2023-05-05 | 4.2e+22 | Confident | MPT-7B | mpt-7b | 93.40 | 85.88 | 97.73 | ||||
| MPT-30B | mpt-30b | MosaicML | United States of America | 2023-06-22 | Open weights (unrestricted) | 2023-06-22 | 1.9e+23 | Confident | MPT-30B | mpt-30b | 99.88 | 96.89 | 103.05 | ||||
| Falcon-7B | falcon-7b | Technology Innovation Institute | United Arab Emirates | 2023-04-24 | Open weights (unrestricted) | 2023-04-24 | 6.3e+22 | Confident | The 7 billion-parameter model in the Falcon series. | Falcon-7B | falcon-7b | 93.91 | 88.44 | 98.82 | |||
| Falcon-40B | falcon-40b | Technology Innovation Institute | United Arab Emirates | 2023-03-15 | Open weights (unrestricted) | 2023-03-15 | 2.4e+23 | Confident | The 40 billion-parameter model in the Falcon series. | Falcon-40B | falcon-40b | 103.53 | 100.09 | 108.57 | |||
| LLaMA-7B | LLaMA-7B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-24 | 4.0e+22 | Confident | Meta's smallest model in the original Llama series. | LLaMA-7B | llama-7b | 95.63 | 88.62 | 100.27 | |||
| LLaMA-13B | LLaMA-13B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-27 | 7.8e+22 | Confident | Meta's 13 billion-parameter model in the original Llama series. | LLaMA-13B | llama-13b | 99.60 | 95.40 | 103.07 | |||
| LLaMA-33B | LLaMA-33B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-27 | 2.7e+23 | Confident | Meta's 33 billion-parameter model in the original Llama series | LLaMA-33B | llama-33b | 106.56 | 104.12 | 109.65 | |||
| LLaMA-65B | LLaMA-65B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-24 | 5.5e+23 | Confident | Meta's largest model in the original Llama series. | LLaMA-65B | llama-65b | 109.32 | 106.50 | 112.51 | |||
| Llama 2-34B | Llama-2-34b | Meta AI | United States of America | 2023-07-18 | Unreleased | 2023-07-18 | 4.1e+23 | Confident | Meta's 34 billion-parameter model in the Llama 2 series. | Llama 2-34B | llama-2-34b | 104.33 | 101.64 | 107.91 | |||
| Chinchilla | Chinchilla (70B) | DeepMind | United Kingdom of Great Britain and Northern Ireland | 2022-03-29 | Unreleased | 2022-03-29 | 5.8e+23 | Confident | Chinchilla | Chinchilla | chinchilla | ||||||
| Claude Instant | claude-instant-1.2 | Anthropic | United States of America | 2023-08-09 | API access | 2023-08-09 | Unknown | Claude Instant | claude-instant | 120.53 | 117.29 | 125.18 | |||||
| Claude Opus 4.7 | claude-opus-4-7_unknown | Claude Opus 4.7 (unknown) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (unknown settings) | Claude Opus 4.7 | claude-opus-4-7 | 156.18 | 153.09 | 164.09 | |||
| Claude Opus 4.7 | claude-opus-4-7 | Claude Opus 4.7 (no thinking) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (no thinking) | Claude Opus 4.7 | claude-opus-4-7 | 156.18 | 153.09 | 164.09 | |||
| Claude Opus 4.6 | claude-opus-4-6_unknown | Claude Opus 4.6 (unknown thinking) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (unknown settings) | Claude Opus 4.6 | claude-opus-4-6 | 155.28 | 152.40 | 163.08 | |||
| Qwen3.7-Max | qwen3.7-max | Alibaba | China | 2026-05-19 | API access | 2026-05-19 | Unverified | Qwen3.7-Max | qwen3-7-max | ||||||||
| Gemini 3.5 Flash | gemini-3.5-flash_unknown | Gemini 3.5 Flash (unknown thinking) | United States of America | 2026-05-19 | Unverified | Gemini 3.5 Flash | gemini-3-5-flash | 156.31 | 154.32 | 164.67 | |||||||
| GPT-5.5 | gpt-5.5_xhigh | GPT-5.5 (xhigh) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.22 | 155.43 | 165.84 | ||||
| GPT-5.5 | gpt-5.5_high | GPT-5.5 (high) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.22 | 155.43 | 165.84 | ||||
| DeepSeek-V4-Pro | deepseek-v4-pro_unknown | DeepSeek v4 Pro (unknown thinking) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 9.7e+24 | Likely | DeepSeek v4 (high) | DeepSeek-V4-Pro | deepseek-v4-pro | |||||
| GPT-5.5 | gpt-5.5_unknown | GPT-5.5 (unknown thinking) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.22 | 155.43 | 165.84 | ||||
| GPT-5.4 | gpt-5.4-2026-03-05_medium | GPT-5.4 (medium) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (medium) | GPT-5.4 | gpt-5-4 | 156.13 | 153.44 | 163.96 | |||
| Kimi K2.5 | kimi-k2.5_none | Kimi K2.5 (instant) | Moonshot | China | 2026-01-27 | Open weights (unrestricted) | 2026-02-02 | 5.8e+24 | Likely | Kimi K2.5 | kimi-k2-5 | 148.22 | 145.05 | 154.59 | |||
| GPT-5.3 Codex | gpt-5.3-codex | GPT-5.3 Codex | OpenAI | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | GPT-5.3 Codex | GPT-5.3 Codex | gpt-5-3-codex | 155.89 | 153.05 | 164.24 | |||
| GPT-5.2 | gpt-5.2-2025-12-11_unknown | GPT-5.2 (unknown thinking) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (high) | GPT-5.2 | gpt-5-2 | 153.72 | 150.93 | 161.15 | |||
| MiniMax-M2.7 | MiniMax-M2.7 | MiniMax | China | 2026-03-18 | Open weights (non-commercial) | 2026-03-18 | Unverified | MiniMax-M2.7 | minimax-m2-7 | ||||||||
| Grok 4.20 | grok-4-20 | xAI | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Grok 4.20 | grok-4-20 | 154.28 | 147.97 | 160.74 | |||||
| MiniMax-M2.1 | MiniMax-M2.1 | MiniMax | China | 2025-12-23 | Open weights (restricted use) | 2025-12-23 | Confident | MiniMax-M2.1 | minimax-m2-1 | ||||||||
| GPT-5.4 | gpt-5.4-2026-03-05_unknown | GPT-5.4 | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (none) | GPT-5.4 | gpt-5-4 | 156.13 | 153.44 | 163.96 | |||
| Claude Opus 4.1 | claude-opus-4-1-20250805_unknown | Claude Opus 4.1 (unknown thinking) | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model, benchmarked with up to 27,000 reasoning tokens. | Claude Opus 4.1 (unknown settings) | Claude Opus 4.1 | claude-opus-4-1 | 144.73 | 141.67 | 150.99 | ||
| MiniMax-M2.5 | MiniMax-M2.5 | MiniMax | China | 2026-02-12 | Open weights (unrestricted) | 2026-02-12 | Likely | MiniMax-M2.5 | minimax-m2-5 | 147.44 | 143.57 | 154.27 | |||||
| Grok 4.3 Beta | grok-4-3 | xAI | United States of America | 2026-04-17 | Hosted access (no API) | 2026-04-17 | Likely | Grok 4.3 Beta | grok-4-3-beta | ||||||||
| DeepSeek-V3.2-Exp | DeepSeek-V3.2-Exp_thinking | DeepSeek | China | 2025-09-29 | Open weights (unrestricted) | 2025-09-29 | 4.2e+24 | Confident | DeepSeek-V3.2-Exp | deepseek-v3-2-exp | 145.19 | 141.83 | 151.64 | ||||
| GLM-4.6 | glm-4.6 | GLM-4.6 | Z.ai (Zhipu AI),Tsinghua University | China | 2025-09-30 | Open weights (unrestricted) | 2025-09-30 | 4.4e+24 | Likely | GLM-4.6 | glm-4-6 | 141.38 | 132.95 | 147.00 | |||
| GPT-5.1 | gpt-5.1-2025-11-13_unknown | GPT-5.1 (unknown thinking) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (high) | GPT-5.1 | gpt-5-1 | 149.71 | 146.98 | 156.51 | |||
| GPT-5.2 Codex | gpt-5.2-codex | GPT-5.2 Codex | OpenAI | United States of America | 2025-12-18 | API access | 2025-12-18 | Likely | GPT-5.2 Codex | GPT-5.2 Codex | gpt-5-2-codex | ||||||
| GPT-5.1-Codex | gpt-5.1-codex | GPT-5.1 Codex | OpenAI | United States of America | 2025-11-12 | API access | 2025-11-12 | Unknown | GPT-5.1-codex | GPT-5.1-Codex | gpt-5-1-codex | ||||||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_unknown | Claude Haiku 4.5 (unknown thinking) | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (16k thinking) | Claude Haiku 4.5 | claude-haiku-4-5 | 143.09 | 140.17 | 149.33 | |||
| MiniMax-M2 | MiniMax-M2 | MiniMax | China | 2025-10-27 | Open weights (unrestricted) | 2025-10-27 | Confident | MiniMax-M2 | minimax-m2 | ||||||||
| DeepSeek-V3.2 | deepseek/deepseek-v3.2 | DeepSeek-V3.2 (Thinking; Novita) | DeepSeek | China | 2025-12-01 | Open weights (unrestricted) | 2025-12-01 | 4.2e+24 | Likely | DeepSeek-V3.2 | DeepSeek-V3.2 | deepseek-v3-2 | 146.46 | 143.27 | 152.46 | ||
| Qwen3-Coder-480B-A35B | Qwen3-Coder-480B-A35B-Instruct | Alibaba | China | 2025-07-31 | Open weights (unrestricted) | 2025-07-22 | 1.6e+24 | Confident | Qwen3 Coder | Qwen3-Coder-480B-A35B | qwen3-coder-480b-a35b | ||||||
| Gemini 3.1 Flash-Lite | gemini-3.1-flash-lite | United States of America | 2026-03-03 | API access | 2026-03-03 | Likely | Gemini 3.1 Flash-Lite | gemini-3-1-flash-lite | |||||||||
| GPT-5.1-Codex-mini | gpt-5.1-codex-mini | GPT-5.1 Codex Mini | OpenAI | United States of America | 2025-11-12 | API access | 2025-11-12 | Unknown | GPT-5.1-codex mini | GPT-5.1-Codex-mini | gpt-5-1-codex-mini | ||||||
| Grok 4.1 Fast | grok-4-1-fast-reasoning | xAI | United States of America | 2025-11-19 | API access | 2025-11-19 | Unknown | Grok 4.1 Fast | grok-4-1-fast | ||||||||
| Grok 4.1 | grok-4-1 | xAI | United States of America | 2025-11-17 | API access | 2025-11-17 | Unknown | Grok 4.1 | grok-4-1 | ||||||||
| Grok 4 Fast | grok-4-fast | xAI | United States of America | 2025-09-19 | API access | 2025-09-19 | Unknown | Grok 4 Fast | Grok 4 Fast | grok-4-fast | 144.89 | 141.05 | 153.40 | ||||
| grok-code-fast-1 | 2025-08-28 | Unverified | grok-code-fast-1 | ||||||||||||||
| QwQ-32B | QwQ-32B | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B | QwQ-32B | QwQ-32B | qwq-32b | |||||
| Gemini 2.0 Flash | gemini-exp-1206 | Google DeepMind,Google | United States of America | 2024-12-06 | API access | 2024-12-11 | Unknown | A December 2024 experimental version of Gemini 2.0 Flash, a small langauge model from Google DeepMind. | Gemini 2.0 Flash (Dec 2024) | Gemini 2.0 Flash (Dec 2024) | gemini-2-0-flash-dec-2024 | 135.31 | 125.82 | 141.06 | |||
| o3-mini | o3-mini-2025-01-31_low | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | Unknown | A smaller version of OpenAI's o3 reasoning model, evaluated with low reasoning effort. | o3-mini | o3-mini | 141.63 | 138.79 | 147.93 | ||||
| Qwen2.5-Max | qwen2.5-max | Alibaba | China | 2025-01-28 | API access | 2025-01-28 | Unknown | The largest model in the Qwen 2.5 series. | Qwen2.5-Max | Qwen2.5-Max | qwen2-5-max | 133.49 | 129.89 | 139.94 | |||
| Gemini 2.0 Flash | gemini-2.0-flash-exp | Google DeepMind,Google | United States of America | 2024-12-11 | API access | 2024-12-11 | Unknown | A December 2024 experimental version of Gemini 2.0 Flash, a small langauge model from Google DeepMind. | Gemini 2.0 Flash (Dec 2024) | Gemini 2.0 Flash (Dec 2024) | gemini-2-0-flash-dec-2024 | 135.31 | 125.82 | 141.06 | |||
| Gemini 2.0 Flash-Lite | gemini-2.0-flash-lite | Google DeepMind | United States of America | 2025-02-05 | API access | 2024-02-05 | Unknown | The smallest model in Google DeepMind's Gemini 2.0 series. | Gemini 2.0 Flash-Lite | gemini-2-0-flash-lite | |||||||
| Gemini 2.0 Flash-Lite | gemini-2.0-flash-lite-preview-02-05 | Google DeepMind | United States of America | 2025-02-05 | API access | 2024-02-05 | Unknown | A February 2025 preview version of the smallest model in Google DeepMind's Gemini 2.0 series. | Gemini 2.0 Flash-Lite | gemini-2-0-flash-lite | |||||||
| Dracarys2-72B-Instruct | 2024-09-30 | Unverified | dracarys2-72b-instruct | ||||||||||||||
| learnlm-1.5-pro-experimental | Unverified | learnlm-1-5-pro-experimental | |||||||||||||||
| Qwen2.5-Coder (32B) | Qwen2.5-Coder-32B-Instruct | Alibaba | China | 2024-11-21 | Open weights (unrestricted) | 2024-11-12 | 1.1e+24 | Confident | An instruction-tuned and coding-optimized 32 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-Coder-32B (instruct) | Qwen2.5-Coder-32B-Instruct | Qwen2.5-Coder-32B-Instruct | qwen2-5-coder-32b-instruct | ||||
| Dracarys2-Llama-3.1-70B-Instruct | 2024-08-14 | Unverified | dracarys2-llama-3-1-70b-instruct | ||||||||||||||
| DeepSeek-R1-Distill-Qwen-32B | DeepSeek-R1-Distill-Qwen-32B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | Speculative | A model based on Qwen 32B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | DeepSeek-R1-Distill-Qwen-32B | deepseek-r1-distill-qwen-32b | |||||||
| QwQ-32B | QwQ-32B-Preview | Alibaba | China | 2024-11-28 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B Preview | QwQ-32B-Preview | QwQ-32B-Preview | qwq-32b-preview | |||||
| Amazon Nova Lite | amazon.nova-lite-v1:0 | Amazon | United States of America | 2024-12-03 | API access | 2024-12-03 | Unknown | Amazon Nova Lite | amazon-nova-lite | ||||||||
| Command R+ | c4ai-command-r-plus-08-2024 | Cohere,Cohere for AI | Canada | 2024-08-30 | Open weights (non-commercial) | 2024-04-04 | Confident | Command R+ | Command R+ | command-r | |||||||
| Amazon Nova Micro | amazon.nova-micro-v1:0 | Amazon | United States of America | 2024-12-03 | API access | 2024-12-03 | Unknown | Amazon Nova Micro | amazon-nova-micro | ||||||||
| c4ai-command-r-08-2024 | 2024-08-30 | Unverified | c4ai-command-r-08-2024 | ||||||||||||||
| OLMo 2 Furious 13B | OLMo-2-1124-13B-Instruct | Allen Institute for AI,University of Washington,New York University (NYU) | United States of America | 2024-12-31 | Open weights (unrestricted) | 2024-12-31 | 4.6e+23 | Confident | OLMo 2 Furious 13B | olmo-2-furious-13b | |||||||
| computer-use-preview-2025-03-11 | 2025-03-11 | Unverified | computer-use-preview-2025-03-11 | ||||||||||||||
| Qwen2.5-72B | Qwen2.5-VL-72B-Instruct | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | Confident | A 72 billion-parameter instruction-tuned vision language model in the Qwen 2.5 series. | Qwen2.5-VL-72B | Qwen2.5-72B | qwen2-5-72b | 129.40 | 124.88 | 133.16 | ||
| gpt-oss-120b | gpt-oss-120b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | The larger one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-120b | gpt-oss-120b | 140.79 | 135.32 | 148.00 | |||
| Qwen3-235B-A22B | Qwen3-235B-A22B | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Likely | Qwen3 (Apr 2025) | Qwen3-235B-A22B | qwen3-235b-a22b | 139.73 | 135.99 | 145.21 | |||
| openhands-lm-32b-v0.1 | 2024-03-26 | Unverified | openhands-lm-32b-v0-1 | ||||||||||||||
| Kimi K2 | Kimi-K2-Instruct | Kimi K2 Instruct | Moonshot | China | 2025-07-12 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | Moonshot AI’s 1-trillion parameter scale model. | Kimi K2 (Jul 2025) | Kimi K2 (Jul 2025) | kimi-k2-jul-2025 | 140.60 | 137.19 | 147.28 | |
| o3-pro | o3-pro-2025-06-10_high | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | Unknown | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with high reasoning effort. | o3-pro | o3-pro | 148.13 | 144.61 | 155.84 | |||||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro-preview-06-05_32K | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America | 2025-06-05 | API access | 2025-06-05 | Unknown | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model, evaluated with up to 32,000 reasoning tokens. | Gemini 2.5 Pro Preview (Jun 2025, 32k thinking) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.72 | 143.80 | 153.41 | |
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_23K | Google DeepMind | United States of America | 2025-05-20 | API access | 2025-04-17 | Unknown | A May 2025 preview version of Gemini 2.5 Flash, a small model in Google DeepMind's Gemini 2.5 series, evaluated with up to 23,000 reasoning tokens. | Gemini 2.5 Flash (May 2025) | Gemini 2.5 Flash (May 2025) | gemini-2-5-flash-may-2025 | 142.45 | 140.05 | 148.50 | |||
| Claude Opus 4 | claude-opus-4-20250514_32K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 32,000 reasoning tokens. | Claude Opus 4 (32k thinking) | Claude Opus 4 | claude-opus-4 | 143.37 | 140.42 | 150.37 | ||
| Qwen3-32B | Qwen3-32B | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 7.1e+24 | Confident | The 32 billion-parameter model in the Qwen 3 series. | Qwen3 32B | Qwen3-32B | qwen3-32b | |||||
| Llama 4 Maverick | Llama-4-Maverick-17B-128E-Instruct | Llama 4 Maverick | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | Likely | The 17 billion active (400 billion total)-parameter model in Meta's Llama 4 series. | Llama 4 Maverick | llama-4-maverick | 133.13 | 129.73 | 138.87 | ||
| GPT-4o | chatgpt-4o-03-27-2025 | OpenAI | United States of America | 2025-03-27 | API access | 2024-05-13 | Speculative | A version of GPT-4o that was behind the ChatGPT interface, released in March 2025. | ChatGPT-4o (Mar 2025) | GPT-4o (Mar 2025) | GPT-4o (Mar 2025) | gpt-4o-mar-2025 | |||||
| Cohere Command A | c4ai-command-a-03-2025 | Cohere | Canada | 2025-03-13 | Open weights (non-commercial) | 2025-03-13 | Confident | Command A | Cohere Command A | cohere-command-a | |||||||
| DeepSeek-V2.5 | DeepSeek-V2.5 | DeepSeek | China | 2024-09-06 | Open weights (restricted use) | 2024-09-06 | 1.8e+24 | Confident | DeepSeek-V2.5 (Sept 2024) | DeepSeek-V2.5 (Sep 2024) | DeepSeek-V2.5 (Sep 2024) | deepseek-v2-5-sep-2024 | |||||
| GPT-4o | chatgpt-4o-01-29-2025 | OpenAI | United States of America | 2025-01-29 | API access | 2024-05-13 | Speculative | A version of GPT-4o that was behind the ChatGPT interface, released in January 2025. | ChatGPT-4o (Jan 2025) | GPT-4o (Jan 2025) | GPT-4o (Jan 2025) | gpt-4o-jan-2025 | |||||
| Codestral | codestral-2501 | Mistral AI | France | 2025-01-13 | Open weights (non-commercial) | 2024-05-29 | Confident | A January 2025 version of a 2024 coding-optimized model from Mistral. | Codestral | codestral | |||||||
| Yi-Lightning | yi-lightning | 01.AI | China | 2024-12-02 | API access | 2024-10-18 | 1.5e+24 | Confident | Yi-Lightning | Yi-Lightning | yi-lightning | ||||||
| o1-mini | o1-mini-2024-09-12_unknown | o1-mini (unknown thinking) | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | Unknown | A smaller version of OpenAI’s first reasoning model, o1, evaluated with high reasoning effort. | o1-mini | o1-mini | 136.68 | 133.18 | 142.63 | |||
| Qwen3-235B-A22B (Jul 2025) | Qwen3-235B-A22B-Instruct-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | A 22 billion active (235 billion total)-parameter mixture-of-experts instruction-tuned model in the Qwen 3 series. | Qwen3 Non-thinking (Jul 2025) | Qwen3-235B-A22B-Instruct (Jul 2025) | Qwen3-235B-A22B-Instruct (Jul 2025) | qwen3-235b-a22b-instruct-jul-2025 | 139.11 | 135.96 | 144.57 | |
| o3 | o3-2025-04-16_unknown | o3 (unknown thinking) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | Unknown | OpenAI's second generation o-series reasoning model, evaluated with high reasoning effort. | o3 | o3 | 147.27 | 144.19 | 154.26 | |||
| Grok 4 | grok-4-0709_high | xAI | United States of America | 2025-07-09 | API access | 2025-07-09 | 5.0e+26 | Speculative | Grok 4 | Grok 4 | grok-4 | 147.43 | 144.38 | 154.25 | |||
| Kimi K2 | moonshotai/kimi-k2-0905 | Kimi K2 0905 (Novita) | Moonshot | China | 2025-09-05 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | A July 2025 preview version of Kimi K2 turbo, an updated version of Kimi K2. | Kimi K2 (Sep 2025) | Kimi K2 (Sep 2025) | kimi-k2-sep-2025 | 141.16 | 138.25 | 147.42 | |
| DeepSeek-V3.2 | deepseek-chat | DeepSeek | China | 2025-12-01 | Open weights (unrestricted) | 2025-12-01 | 4.2e+24 | Likely | DeepSeek-V3.2 | deepseek-v3-2 | 146.46 | 143.27 | 152.46 | ||||
| Claude Opus 4.5 | claude-opus-4-5-20251101_unknown | Claude Opus 4.5 (unknown thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (unknown settings) | Claude Opus 4.5 | claude-opus-4-5 | 149.89 | 147.32 | 157.25 | |||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_unknown | Claude Sonnet 4.5 (unknown thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (59k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 147.15 | 144.42 | 153.64 | |||
| Mixtral 8x22B | Mixtral-8x22B-Instruct-v0.1 | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | Speculative | An instruction-tuned version of Mistral's 8 billion active (176 billion total)-parameter mixture-of-experts model | Mixtral 8x22B | mixtral-8x22b | 121.21 | 114.22 | 125.22 | |||
| Gemini 1.5 Pro | gemini-1.5-pro-001-feb24 | Google DeepMind | United States of America | 2024-02-15 | API access | 2024-02-15 | Speculative | The first of two versions of a previous flagship model from Google DeepMind, Gemini 1.5 Pro. | Gemini 1.5 Pro (Feb 2024) | Gemini 1.5 Pro (Feb 2024) | gemini-1-5-pro-feb-2024 | 127.78 | 125.40 | 136.15 | |||
| GPT-5.1-Codex-Max | gpt-5.1-codex-max | GPT-5.1-Codex-Max | OpenAI | United States of America | 2025-11-19 | API access | 2025-11-19 | Unknown | GPT-5.1-Codex-Max | GPT-5.1-Codex-Max | gpt-5-1-codex-max | ||||||
| Claude Opus 4.5 | claude-opus-4-5-20251101_128K | Claude Opus 4.5 (128k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (128k thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.89 | 147.32 | 157.25 | |||
| Claude Sonnet 4.6 | claude-sonnet-4-6_unknown | Claude Sonnet 4.6 (unknown thinking) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 152.58 | 148.95 | 158.53 | ||||
| GPT‑5-Codex | gpt-5-codex | GPT-5-codex | OpenAI | United States of America | 2025-09-15 | API access | 2025-09-15 | Unknown | GPT-5-codex | GPT‑5-Codex | gpt5-codex | ||||||
| Kimi K2 Thinking | kimi-k2-thinking | Kimi K2 Thinking | Moonshot | China | 2025-11-06 | Open weights (restricted use) | 2025-11-06 | 4.2e+24 | Likely | Kimi K2 Thinking | kimi-k2-thinking | 145.60 | 142.44 | 152.09 | |||
| GPT-5 mini | gpt-5-mini-2025-08-07_unknown | GPT-5 mini (unknown thinking) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 mini (high) | GPT-5 mini | gpt-5-mini | 145.59 | 142.41 | 152.17 | ||
| Qwen 3.6 35B-A3B | qwen3.6-35b-a3b | Alibaba | China | 2026-04-14 | Open weights (unrestricted) | 2026-04-14 | Unverified | Qwen 3.6 35B-A3B | qwen-3-6-35b-a3b | ||||||||
| GPT-5 nano | gpt-5-nano-2025-08-07_unknown | GPT-5 nano (unknown thinking) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 nano (high) | GPT-5 nano | gpt-5-nano | 140.95 | 137.13 | 147.50 | ||
| gpt-oss-120b | gpt-oss-120b | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | The larger one of OpenAI's two latest open source language models. | gpt-oss-120b | gpt-oss-120b | 140.79 | 135.32 | 148.00 | |||
| gpt-oss-120b | gpt-oss-120b_unknown | gpt-oss-120b (unknown thinking) | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | The larger one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-120b | gpt-oss-120b | 140.79 | 135.32 | 148.00 | ||
| gpt-oss-20b | gpt-oss-20b | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | Confident | The smaller one of OpenAI's two latest open source language models. | gpt-oss-20b | gpt-oss-20b | ||||||
| gpt-oss-20b | gpt-oss-20b_unknown | gpt-oss-20b (unknown thinking) | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | Confident | The smaller one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-20b | gpt-oss-20b | |||||
| QwQ-32B | QwQ-32B (16K thinking) | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B (16k thinking) | QwQ-32B | QwQ-32B | qwq-32b | |||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-04-17 (24K thinking) | Google DeepMind | United States of America | 2025-04-17 | API access | 2025-04-17 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 24,000 reasoning tokens. | Gemini 2.5 Flash (Apr 2025) | Gemini 2.5 Flash (Apr 2025) | gemini-2-5-flash-apr-2025 | 140.71 | 137.20 | 146.83 | |||
| Qwen3-30B-A3B | Qwen3-30B-A3B | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 6.5e+23 | Likely | Qwen3-30B-A3B | Qwen3-30B-A3B | qwen3-30b-a3b | ||||||
| o3-pro | o3-pro-2025-06-10_medium | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | Unknown | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with medium reasoning effort. | o3-pro | o3-pro | 148.13 | 144.61 | 155.84 | |||||
| DeepSeek-V3.1 | DeepSeek-V3.1_thinking | DeepSeek | China | 2025-08-21 | Open weights (unrestricted) | 2025-08-21 | 3.6e+24 | Confident | DeepSeek-V3.1 (thinking) | DeepSeek-V3.1 | deepseek-v3-1 | 138.89 | 136.11 | 150.69 | |||
| DeepSeek-V3.1 | DeepSeek-V3.1 | DeepSeek | China | 2025-08-21 | Open weights (unrestricted) | 2025-08-21 | 3.6e+24 | Confident | An updated version of DeepSeek V3. | DeepSeek-V3.1 | DeepSeek-V3.1 | deepseek-v3-1 | 138.89 | 136.11 | 150.69 | ||
| gemini-3-deep-think-preview | 2026-02-12 | Unverified | gemini-3-deep-think-preview | ||||||||||||||
| GPT-5.5 Pro | gpt-5.5-pro_high | GPT-5.5 Pro (high) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 Pro | gpt-5-5-pro | 159.35 | 155.79 | 167.47 | ||||
| GPT-5.5 Pro | gpt-5.5-pro_xhigh | GPT-5.5 Pro (xhigh) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 Pro | gpt-5-5-pro | 159.35 | 155.79 | 167.47 | ||||
| GPT-5.5 | gpt-5.5_medium | GPT-5.5 (medium) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.22 | 155.43 | 165.84 | ||||
| Claude Opus 4.7 | claude-opus-4-7_high | Claude Opus 4.7 (high) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (high) | Claude Opus 4.7 | claude-opus-4-7 | 156.18 | 153.09 | 164.09 | |||
| Claude Opus 4.7 | claude-opus-4-7_low | Claude Opus 4.7 (low) | Anthropic | United States of America | 2026-04-16 | API access | 2026-04-16 | Likely | Claude Opus 4.7 (low) | Claude Opus 4.7 | claude-opus-4-7 | 156.18 | 153.09 | 164.09 | |||
| Claude Sonnet 4.6 | claude-sonnet-4-6_high | Claude Sonnet 4.6 (high) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 152.58 | 148.95 | 158.53 | ||||
| Claude Sonnet 4.6 | claude-sonnet-4-6_max | Claude Sonnet 4.6 (max) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 152.58 | 148.95 | 158.53 | ||||
| GPT-5.2 Pro | gpt-5.2-pro-2025-12-11_high | GPT-5.2 Pro | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Likely | GPT-5.2 Pro (high) | GPT-5.2 Pro | gpt-5-2-pro | 154.19 | 150.15 | 162.75 | |||
| GPT-5.2 Pro | gpt-5.2-pro-2025-12-11_medium | GPT-5.2 Pro | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Likely | GPT-5.2 Pro (medium) | GPT-5.2 Pro | gpt-5-2-pro | 154.19 | 150.15 | 162.75 | |||
| Claude Opus 4.5 | claude-opus-4-5-20251101_64K | Claude Opus 4.5 (64k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (64k thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.89 | 147.32 | 157.25 | |||
| GPT-5.4 | gpt-5.4-2026-03-05_low | GPT-5.4 (low) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (low) | GPT-5.4 | gpt-5-4 | 156.13 | 153.44 | 163.96 | |||
| GPT-5.4 Mini | gpt-5.4-mini-2026-03-17_xhigh | GPT-5.4 mini (xhigh) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 mini (high) | GPT-5.4 Mini | gpt-5-4-mini | 148.74 | 145.69 | 155.92 | |||
| GPT-5 Pro | gpt-5-pro-2025-10-06_unknown | GPT-5 Pro (unknown thinking) | OpenAI | United States of America | 2025-10-07 | API access | 2025-10-07 | Unknown | GPT-5 Pro | gpt-5-pro | 150.21 | 146.92 | 157.69 | ||||
| Claude Opus 4.5 | claude-opus-4-5-20251101_8K | Claude Opus 4.5 (8k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (8k thinking) | Claude Opus 4.5 | claude-opus-4-5 | 149.89 | 147.32 | 157.25 | |||
| Gemini 3.5 Flash | gemini-3.5-flash_minimal | Gemini 3.5 Flash (minimal) | United States of America | 2026-05-19 | Unverified | Gemini 3.5 Flash | gemini-3-5-flash | 156.31 | 154.32 | 164.67 | |||||||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_8K | Claude Sonnet 4.5 (8k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (2k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 147.15 | 144.42 | 153.64 | |||
| tiny-recursion-model | 2025-10-06 | Unverified | tiny-recursion-model | ||||||||||||||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_1K | Claude Sonnet 4.5 (1k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (2k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 147.15 | 144.42 | 153.64 | |||
| GPT-5.4 Nano | gpt-5.4-nano-2026-03-17_xhigh | GPT-5.4 nano (xhigh) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 nano (high) | GPT-5.4 Nano | gpt-5-4-nano | 146.23 | 142.65 | 153.80 | |||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro_32K | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-05 | Unknown | Google’s general-purpose frontier reasoning model, evaluated with up to 32,000 reasoning tokens. | Gemini 2.5 Pro (32k thinking) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.72 | 143.80 | 153.41 | ||
| Claude Opus 4 | claude-opus-4-20250514_8K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 8,000 reasoning tokens. | Claude Opus 4 (8k thinking) | Claude Opus 4 | claude-opus-4 | 143.37 | 140.42 | 150.37 | ||
| GPT-5.4 Mini | gpt-5.4-mini-2026-03-17_medium | GPT-5.4 mini (medium) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 mini (high) | GPT-5.4 Mini | gpt-5-4-mini | 148.74 | 145.69 | 155.92 | |||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro_16K | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-05 | Unknown | Google’s general-purpose frontier reasoning model, evaluated with up to 16,000 reasoning tokens. | Gemini 2.5 Pro (16k thinking) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.72 | 143.80 | 153.41 | ||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro_8K | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-05 | Unknown | Google’s general-purpose frontier reasoning model, evaluated with up to 8,000 reasoning tokens. | Gemini 2.5 Pro (8k thinking) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.72 | 143.80 | 153.41 | ||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_16K | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (16k thinking) | Claude Haiku 4.5 | claude-haiku-4-5 | 143.09 | 140.17 | 149.33 | ||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_1K | Google DeepMind | United States of America | 2025-05-20 | API access | 2025-04-17 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 1,000 reasoning tokens. | Gemini 2.5 Flash (May 2025) | Gemini 2.5 Flash (May 2025) | gemini-2-5-flash-may-2025 | 142.45 | 140.05 | 148.50 | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_8K | Google DeepMind | United States of America | 2025-05-20 | API access | 2025-04-17 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 8,000 reasoning tokens. | Gemini 2.5 Flash (May 2025) | Gemini 2.5 Flash (May 2025) | gemini-2-5-flash-may-2025 | 142.45 | 140.05 | 148.50 | |||
| Claude Sonnet 4 | claude-sonnet-4-20250514_8K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Claude Sonnet 4 (8k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.57 | 138.67 | 149.90 | ||||
| o3-pro | o3-pro-2025-06-10_low | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | Unknown | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with low reasoning effort. | o3-pro | o3-pro | 148.13 | 144.61 | 155.84 | |||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_16K | Google DeepMind | United States of America | 2025-05-20 | API access | 2025-04-17 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 16,000 reasoning tokens. | Gemini 2.5 Flash (May 2025) | Gemini 2.5 Flash (May 2025) | gemini-2-5-flash-may-2025 | 142.45 | 140.05 | 148.50 | |||
| GPT-5.4 Nano | gpt-5.4-nano-2026-03-17_medium | GPT-5.4 nano (medium) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 nano (high) | GPT-5.4 Nano | gpt-5-4-nano | 146.23 | 142.65 | 153.80 | |||
| GPT-5 mini | gpt-5-mini-2025-08-07_minimal | GPT-5 mini (minimal) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 mini (medium) | GPT-5 mini | gpt-5-mini | 145.59 | 142.41 | 152.17 | ||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_8K | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (8k thinking) | Claude Haiku 4.5 | claude-haiku-4-5 | 143.09 | 140.17 | 149.33 | ||||
| GPT-5.4 Nano | gpt-5.4-nano-2026-03-17_low | GPT-5.4 nano (low) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 nano (high) | GPT-5.4 Nano | gpt-5-4-nano | 146.23 | 142.65 | 153.80 | |||
| codex-mini-2025-05-16 | 2025-05-16 | Unverified | codex-mini-2025-05-16 | ||||||||||||||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_1K | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (1k thinking) | Claude Haiku 4.5 | claude-haiku-4-5 | 143.09 | 140.17 | 149.33 | ||||
| GPT-5.4 Mini | gpt-5.4-mini-2026-03-17_low | GPT-5.4 mini (low) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 mini (high) | GPT-5.4 Mini | gpt-5-4-mini | 148.74 | 145.69 | 155.92 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_8K | Claude 3.7 Sonnet (8k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 9,000 reasoning tokens allowed. | Claude 3.7 Sonnet (8k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 142.03 | 139.52 | 148.45 | |
| Claude Sonnet 4 | claude-sonnet-4-20250514_1K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized Claude 4 series model, evaluated at up to 1,000 reasoning tokens. | Claude Sonnet 4 (1k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.57 | 138.67 | 149.90 | |||
| GPT-5 mini | gpt-5-mini-2025-08-07_low | GPT-5 mini (low) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 mini (medium) | GPT-5 mini | gpt-5-mini | 145.59 | 142.41 | 152.17 | ||
| Grok-3 mini | grok-3-mini_low | Grok 3 mini | xAI | United States of America | 2025-06-24 | API access | 2025-02-19 | Unknown | Grok 3 mini (low) | Grok-3 mini | grok-3-mini | 141.19 | 138.81 | 147.60 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_1K | Claude 3.7 Sonnet (1k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 1,000 reasoning tokens allowed. | Claude 3.7 Sonnet (1k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 142.03 | 139.52 | 148.45 | |
| Grok 3 | grok-3 | xAI | United States of America | 2025-04-09 | API access | 2025-02-17 | 3.5e+26 | Likely | XAI's third generation flagship model. | Grok 3 | Grok 3 | grok-3 | 139.18 | 136.31 | 144.98 | ||
| Claude Opus 4 | claude-opus-4-20250514_1K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 1,000 reasoning tokens. | Claude Opus 4 (1k thinking) | Claude Opus 4 | claude-opus-4 | 143.37 | 140.42 | 150.37 | ||
| Magistral Medium 1.1 | magistral-medium-2506 | Mistral AI | France | 2025-06-10 | API access | 2025-06-10 | Unknown | Magistral Medium 1.1 | magistral-medium-1-1 | ||||||||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro-preview-06-05_1K | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America | 2025-06-05 | API access | 2025-06-05 | Unknown | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model, evaluated with up to 1,000 reasoning tokens. | Gemini 2.5 Pro Preview (Jun 2025, 1k thinking) | Gemini 2.5 Pro (Jun 2025) | Gemini 2.5 Pro (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.72 | 143.80 | 153.41 | |
| GPT-5 nano | gpt-5-nano-2025-08-07_low | GPT-5 nano (low) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 nano (low) | GPT-5 nano | gpt-5-nano | 140.95 | 137.13 | 147.50 | ||
| GPT-5 nano | gpt-5-nano-2025-08-07_minimal | GPT-5 nano (minimal) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 nano (low) | GPT-5 nano | gpt-5-nano | 140.95 | 137.13 | 147.50 | ||
| GPT-5.4 Pro | gpt-5.4-pro-2026-03-05_unknown | GPT-5.4 Pro | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 Pro | GPT-5.4 Pro | gpt-5-4-pro | 157.73 | 155.08 | 165.39 | |||
| Claude Opus 4 | claude-opus-4-20250514_unknown | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 32,000 reasoning tokens. | Claude Opus 4 (unknown settings) | Claude Opus 4 | claude-opus-4 | 143.37 | 140.42 | 150.37 | ||
| GLM-4.5-Air | GLM-4.5-Air | Z.ai (Zhipu AI),Tsinghua University | China | 2025-07-20 | Open weights (unrestricted) | 2025-08-05 | 1.7e+24 | Confident | A smaller version of GLM-4.5, Zhipu AI & Tsinghua University's large open source model. | GLM-4.5-Air | glm-4-5-air | ||||||
| o1-pro | o1-pro-2025-03-19 | o1 Pro | OpenAI | United States of America | 2025-03-19 | API access | 2025-03-19 | Likely | An An enhanced version of OpenAI's first reasoning model,o1, evaluated with low reasoning effort. | o1-pro | o1-pro | ||||||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_unknown | Claude 3.7 Sonnet (unknown thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 64,000 reasoning tokens allowed. | Claude 3.7 Sonnet (64k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 142.03 | 139.52 | 148.45 | |
| o1 | o1-2024-12-17_unknown | o1 (unknown thinking) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | Unknown | OpenAI's first reasoning model, evaluated with at the high reasoning effort level. | o1 | o1 | 142.83 | 139.83 | 148.65 | |||
| Claude Sonnet 4 | claude-sonnet-4-20250514_unknown | Claude Sonnet 4 (unknown thinking) | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized Claude 4 series model, evaluated at up to 12,000 reasoning tokens. | Claude Sonnet 4 (12k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.57 | 138.67 | 149.90 | ||
| T5-Base | Unverified | t5-base | |||||||||||||||
| Switch-Base | Unverified | switch-base | |||||||||||||||
| T5-Large | Unverified | t5-large | |||||||||||||||
| Switch-Large | Unverified | switch-large | |||||||||||||||
| T5-Small | Unverified | t5-small | |||||||||||||||
| T5-3B | T5-3B | United States of America | Open weights (unrestricted) | 2019-10-23 | 9.0e+21 | Confident | The 3 billion-parameter version of T5, an early large language model from Google. | T5-3B | t5-3b | ||||||||
| T5-11B | T5-11B | United States of America | Open weights (unrestricted) | 2019-10-23 | 3.3e+22 | Confident | The largest, 11 billion-parameter version of T5, an early large language model from Google. | T5-11B | t5-11b | ||||||||
| GPT-4 Turbo (Nov 2023) | gpt-4-turbo | OpenAI | United States of America | 2023-11-06 | API access | 2023-11-06 | Unknown | An updated version of GPT-4 and OpenAI's then-new flagship language model. | GPT-4 Turbo (Nov 2023) | gpt-4-turbo-nov-2023 | |||||||
| GPT-3.5 Turbo Instruct | gpt-3.5-turbo-instruct | OpenAI | United States of America | 2023-09-18 | API access | 2023-09-28 | Likely | An instruction-tuned version of OpenAI's GPT-3.5 turbo. | GPT-3.5 Turbo Instruct | gpt-3-5-turbo-instruct | |||||||
| GPT-3 175B (davinci) | davinci-002 | OpenAI | United States of America | API access | 2020-05-28 | 3.1e+23 | Confident | GPT-3 175B (davinci-002) | GPT-3 175B (davinci) | gpt-3-175b-davinci | |||||||
| GPT-2 (1.5B) | gpt2-xl | OpenAI | United States of America | 2019-11-05 | Open weights (unrestricted) | 2019-02-14 | 1.9e+21 | Speculative | The largest version of GPT-2, OpenAI's second generation transformer-based large pretrained language model. | GPT-2 (1.5B) | gpt-2-1-5b | ||||||
| gemini-robotics-er-1.5-preview | 2025-09-26 | Unverified | gemini-robotics-er-1-5-preview | ||||||||||||||
| Gemini 2.5 Flash-Lite | gemini-2.5-flash-lite-preview-09-2025 | Google DeepMind | United States of America | 2025-09-25 | API access | 2025-06-15 | Unknown | Gemini 2.5 Flash-Lite (Sep 2025) | Gemini 2.5 Flash-Lite (Sep 2025) | Gemini 2.5 Flash-Lite (Sep 2025) | gemini-2-5-flash-lite-sep-2025 | ||||||
| GPT-5.5 Pro | gpt-5.5-pro_unknown | GPT-5.5 Pro (unknown thinking) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 Pro | gpt-5-5-pro | 159.35 | 155.79 | 167.47 | ||||
| Claude Opus 4 | claude-opus-4-20250514_12K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 12,000 reasoning tokens. | Claude Opus 4 (12k thinking) | Claude Opus 4 | claude-opus-4 | 143.37 | 140.42 | 150.37 | ||
| GPT-5.2 Pro | gpt-5.2-pro-2025-12-11 | GPT-5.2 Pro | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Likely | GPT-5.2 Pro | GPT-5.2 Pro | gpt-5-2-pro | 154.19 | 150.15 | 162.75 | |||
| Grok 4.1 Fast | grok-4-1-fast-non-reasoning | xAI | United States of America | 2025-11-19 | API access | 2025-11-19 | Unknown | Grok 4.1 Fast | grok-4-1-fast | ||||||||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_12K | Claude Sonnet 4.5 (12k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (12k thinking) | Claude Sonnet 4.5 | claude-sonnet-4-5 | 147.15 | 144.42 | 153.64 | |||
| DeepSeek-V3.2 | DeepSeek-V3.2-Speciale | DeepSeek | China | 2025-12-01 | Open weights (unrestricted) | 2025-12-01 | 4.2e+24 | Likely | DeepSeek-V3.2-Speciale | DeepSeek-V3.2-Speciale | deepseek-v3-2-speciale | ||||||
| GLM-4.7 | zai-org/glm-4.7 | GLM-4.7 (Novita) | Z.ai (Zhipu AI) | China | 2025-12-22 | Open weights (unrestricted) | 2025-12-22 | 4.4e+24 | Likely | GLM-4.7 | glm-4-7 | 144.61 | 141.55 | 150.45 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_12K | Claude 3.7 Sonnet (12k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 12,000 reasoning tokens allowed. | Claude 3.7 Sonnet (12k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 142.03 | 139.52 | 148.45 | |
| Claude Sonnet 4 | claude-sonnet-4-20250514_12K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized Claude 4 series model, evaluated at up to 12,000 reasoning tokens. | Claude Sonnet 4 (12k thinking) | Claude Sonnet 4 | claude-sonnet-4 | 142.57 | 138.67 | 149.90 | |||
| video-SALMONN 2+ | video-SALMONN-2plus | ByteDance | China | 2025-06-18 | 2025-06-18 | Likely | video-SALMONN 2+ | video-salmonn-2 | |||||||||
| InternVL2_5-78B | InternVL2_5-78B | Shanghai AI Lab,SenseTime,Tsinghua University,Nanjing University,Fudan University,Chinese University of Hong Kong (CUHK),Shanghai Jiao Tong University | Hong Kong,China | 2024-12-06 | Open weights (restricted use) | 2024-12-06 | Confident | InternVL2_5-78B | internvl2-5-78b | ||||||||
| Qwen2-VL-72B-Instruct | 2024-08-29 | Unverified | qwen2-vl-72b-instruct | ||||||||||||||
| LLaVA-Video-72B-Qwen2 | 2024-09-02 | Unverified | llava-video-72b-qwen2 | ||||||||||||||
| LinVT | 2024-12-06 | Unverified | linvt | ||||||||||||||
| Aria | 2024-09-30 | Unverified | aria | ||||||||||||||
| ViLAMP-llava-qwen | 2025-05-01 | Unverified | vilamp-llava-qwen | ||||||||||||||
| Oryx-1.5-32B | 2024-10-22 | Unverified | oryx-1-5-32b | ||||||||||||||
| LLaVA-OneVision 72B | 2024-08-06 | Unverified | llava-onevision-72b | ||||||||||||||
| VideoLLAMA3-7B | 2024-01-22 | Unverified | videollama3-7b | ||||||||||||||
| LLaVA-Video-7B-Qwen2 | 2024-09-02 | Unverified | llava-video-7b-qwen2 | ||||||||||||||
| LLaVA-Video-7B-Qwen2-TPO | 2025-01-19 | Unverified | llava-video-7b-qwen2-tpo | ||||||||||||||
| VideoChat-Flash-Qwen2-7B_res448 | 2025-01-11 | Unverified | videochat-flash-qwen2-7b-res448 | ||||||||||||||
| ByteVideoLLM-14B | 2024-10-13 | Unverified | bytevideollm-14b | ||||||||||||||
| NVILA 8B | NVILA-8B | NVIDIA,Massachusetts Institute of Technology (MIT),University of California (UC) Berkeley,University of California San Diego,University of Washington,Tsinghua University | China,United States of America | 2024-12-10 | Open weights (non-commercial) | 2024-12-05 | 2.3e+21 | Likely | NVILA 8B | nvila-8b | |||||||
| LiveCC-7B-Instruct | 2024-04-12 | Unverified | livecc-7b-instruct | ||||||||||||||
| Qwen2-VL-7B-Instruct | 2024-08-29 | Unverified | qwen2-vl-7b-instruct | ||||||||||||||
| MiniCPM-o-2_6 | 2025-01-12 | Unverified | minicpm-o-2-6 | ||||||||||||||
| VideoLLAMA2-7B | 2024-06-12 | Unverified | videollama2-7b | ||||||||||||||
| InternVL2-40B | InternVL2-40B | Shanghai AI Lab | China | 2024-07-08 | Open weights (unrestricted) | 2024-07-04 | Confident | InternVL2-40B | internvl2-40b | ||||||||
| MiniCPM-V-2_6 | 2024-08-03 | Unverified | minicpm-v-2-6 | ||||||||||||||
| GPT-4V | gpt-4-1106-vision-preview | OpenAI | United States of America | 2023-11-06 | API access | 2023-09-25 | Unknown | GPT-4V | gpt-4v | ||||||||
| mPLUG-Owl3-7B-241101 | 2024-11-26 | Unverified | mplug-owl3-7b-241101 | ||||||||||||||
| TimeMarker | 2024-11-27 | Unverified | timemarker | ||||||||||||||
| VITA-1.5 | 2024-12-20 | Unverified | vita-1-5 | ||||||||||||||
| kangaroo | 2024-11-13 | Unverified | kangaroo | ||||||||||||||
| VITA | 2024-08-12 | Unverified | vita | ||||||||||||||
| Video-XL-7B | 2024-10-17 | Unverified | video-xl-7b | ||||||||||||||
| Video-CCAM-7B-v1.2 | 2024-09-29 | Unverified | video-ccam-7b-v1-2 | ||||||||||||||
| long-llava-qwen2-7b | 2024-08-30 | Unverified | long-llava-qwen2-7b | ||||||||||||||
| LongVA-7B | 2024-06-13 | Unverified | longva-7b | ||||||||||||||
| Qwen-VL-Max | 2024-01-18 | Unverified | qwen-vl-max | ||||||||||||||
| InternVL-Chat-V1-5 | 2024-04-18 | Unverified | internvl-chat-v1-5 | ||||||||||||||
| SliME-Llama3-8B | 2024-06-02 | Unverified | slime-llama3-8b | ||||||||||||||
| Chat-Uni-Vi-7B-v1.5 + 100k SG-WV | 2024-06-20 | Unverified | chat-uni-vi-7b-v1-5--100k-sg-wv | ||||||||||||||
| Qwen-VL-Chat | 2023-08-20 | Unverified | qwen-vl-chat | ||||||||||||||
| Chat-UniVi-7B-v1.5 | 2024-04-23 | Unverified | chat-univi-7b-v1-5- | ||||||||||||||
| sharegpt4video-8b | 2024-05-27 | Unverified | sharegpt4video-8b | ||||||||||||||
| Video-LLaVA-7B | 2023-11-17 | Unverified | video-llava-7b | ||||||||||||||
| video_chat2_mistral | 2023-11-29 | Unverified | video-chat2-mistral | ||||||||||||||
| ST-LLM | 2024-03-28 | Unverified | st-llm | ||||||||||||||
| Cerebras-GPT-13B | Cerebras-GPT-13B | Cerebras Systems | United States of America | 2023-03-20 | Open weights (unrestricted) | 2023-04-06 | 2.3e+22 | Confident | Cerebras-GPT-13B | cerebras-gpt-13b | 79.86 | 68.56 | 85.64 | ||||
| Dolly 2.0-12b | dolly-v2-12b | Databricks | United States of America | 2023-04-11 | Open weights (unrestricted) | 2023-04-12 | Confident | Dolly 2.0-12b | dolly-2-0-12b | 87.34 | 79.54 | 94.13 | |||||
| Falcon-180B | falcon-180B | Technology Innovation Institute | United Arab Emirates | 2023-09-06 | Open weights (restricted use) | 2023-09-06 | 3.8e+24 | Confident | The 180 billion-parameter model in the Falcon series. | Falcon-180B | falcon-180b | 111.09 | 106.77 | 117.30 | |||
| GPT-J-6B | gpt-j-6b | EleutherAI,LAION | Germany,United States of America | 2021-08-05 | Open weights (unrestricted) | 2021-05-01 | 1.5e+22 | Confident | GPT-J-6B | gpt-j-6b | |||||||
| GPT-Neo-2.7B | gpt-neo-2.7B | EleutherAI | United States of America | 2023-03-30 | Open weights (unrestricted) | 2021-03-21 | 7.9e+21 | Confident | GPT-Neo-2.7B | gpt-neo-2-7b | |||||||
| GPT-NeoX-20B | gpt-neox-20b | EleutherAI | United States of America | 2022-04-07 | Open weights (unrestricted) | 2022-02-09 | 9.3e+22 | Confident | GPT-NeoX-20B | gpt-neox-20b | |||||||
| open_llama_7b | 2023-06-07 | Unverified | open-llama-7b | ||||||||||||||
| OPT-1.3B | opt-1.3b | Meta AI | United States of America | 2022-05-11 | Open weights (non-commercial) | 2022-06-21 | Confident | OPT-1.3B | opt-1-3b | ||||||||
| opt-13b | 2022-05-11 | Unverified | opt-13b | ||||||||||||||
| PaLM 62B | 2022-04-04 | Unverified | palm-62b | ||||||||||||||
| Phi-1.5 | phi-1_5 | Microsoft | United States of America | 2023-09-11 | Open weights (unrestricted) | 2023-09-11 | 1.2e+21 | Confident | Phi-1.5 | phi-1-5 | 90.22 | 59.04 | 99.56 | ||||
| RedPajama-INCITE-7B-Base | 2023-05-04 | Unverified | redpajama-incite-7b-base | ||||||||||||||
| stablelm-tuned-alpha-7b | 2023-04-19 | Unverified | stablelm-tuned-alpha-7b | ||||||||||||||
| vicuna-13b-v1.1 | 2023-04-12 | Unverified | vicuna-13b-v1-1 | ||||||||||||||
| XGen-7B | xgen-7b-8k-base | Salesforce | United States of America | 2023-06-27 | Open weights (unrestricted) | 2023-09-07 | 8.0e+22 | Confident | XGen-7B | XGen-7B | xgen-7b | 91.60 | 86.44 | 95.80 | |||
| GPT-3.5 (davinci-002) | text-davinci-002 | OpenAI | United States of America | 2022-03-15 | API access | 2022-03-15 | 2.6e+24 | Speculative | davinci-002 | davinci-002 | davinci-002 | ||||||
| Llama 3-8B | Meta-Llama-3-8B | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.2e+23 | Confident | Meta's 8 billion-parameter model in the Llama 3 series. | Llama 3-8B | llama-3-8b | 115.92 | 112.53 | 118.72 | |||
| Llama 3-70B | Meta-Llama-3-70B | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.9e+24 | Confident | Meta's 70 billion-parameter model in the Llama 3 series. | Llama 3-70B | llama-3-70b | 122.65 | 120.47 | 126.55 | |||
| Qwen2.5-Coder-0.5B | 2024-09-18 | Unverified | qwen2-5-coder-0-5b | ||||||||||||||
| Qwen2.5-Coder (1.5B) | Qwen2.5-Coder-1.5B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-09-18 | 5.1e+22 | Confident | Qwen's 1.5 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-1.5B | Qwen2.5-Coder (1.5B) | qwen2-5-coder-1-5b | 102.76 | 90.75 | 109.87 | ||
| Qwen2.5-Coder-3B | 2024-09-18 | Unverified | qwen2-5-coder-3b | ||||||||||||||
| Qwen2.5-Coder (7B) | Qwen2.5-Coder-7B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-09-18 | 2.5e+23 | Confident | Qwen's 7 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-7B | Qwen2.5-Coder (7B) | qwen2-5-coder-7b | 112.54 | 104.01 | 118.82 | ||
| Qwen2.5-Coder-14B | 2024-09-18 | Unverified | qwen2-5-coder-14b | ||||||||||||||
| Qwen2.5-Coder (32B) | Qwen2.5-Coder-32B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-11-12 | 1.1e+24 | Confident | Qwen's 32 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-32B | Qwen2.5-Coder-32B | Qwen2.5-Coder-32B | qwen2-5-coder-32b | 118.82 | 112.40 | 126.31 | |
| DeepSeek Coder 1.3B | deepseek-coder-1.3b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 1.6e+22 | Likely | DeepSeek Coder 1.3B (base) | DeepSeek Coder 1.3B | deepseek-coder-1-3b | 62.02 | 55.86 | 73.34 | |||
| StarCoder 2 3B | starcoder2-3b | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | 2024-02-22 | Open weights (restricted use) | 2024-02-29 | 5.9e+22 | Confident | StarCoder 2 3B | StarCoder 2 3B | starcoder-2-3b | 88.57 | 78.65 | 95.15 | |||
| StarCoder 2 7B | starcoder2-7b | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | 2024-02-20 | Open weights (restricted use) | 2024-02-29 | 1.5e+23 | Confident | StarCoder 2 7B | StarCoder 2 7B | starcoder-2-7b | 93.46 | 83.43 | 100.15 | |||
| DeepSeek Coder 6.7B | deepseek-coder-6.7b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 8.0e+22 | Likely | DeepSeek Coder 6.7B (base) | DeepSeek Coder 6.7B | deepseek-coder-6-7b | 89.47 | 81.35 | 96.34 | |||
| DeepSeek-Coder-V2-Lite-Base | 2024-06-13 | Unverified | deepseek-coder-v2-lite-base | ||||||||||||||
| CodeQwen1.5-7B | 2024-04-15 | Unverified | codeqwen1-5-7b | ||||||||||||||
| StarCoder 2 15B | starcoder2-15b | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | 2024-02-20 | Open weights (restricted use) | 2024-02-29 | 3.9e+23 | Confident | StarCoder 2 15B | StarCoder 2 15B | starcoder-2-15b | 104.48 | 96.00 | 111.72 | |||
| DeepSeek Coder 33B | deepseek-coder-33b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 4.0e+23 | Likely | DeepSeek Coder 33B (base) | DeepSeek Coder 33B | deepseek-coder-33b | 96.20 | 88.16 | 102.61 | |||
| DeepSeek-Coder-V2 236B | DeepSeek-Coder-V2-Base | DeepSeek | China | 2024-06-17 | Open weights (restricted use) | 2024-06-17 | 1.3e+24 | Confident | DeepSeek-Coder-V2 (base) | DeepSeek-Coder-V2 236B | deepseek-coder-v2-236b | ||||||
| Yi 6B | Yi-6B | 01.AI | China | 2023-11-02 | Open weights (restricted use) | 2023-11-02 | 1.3e+23 | Confident | Yi-6B (chat) | Yi 6B | yi-6b | 104.06 | 98.93 | 108.63 | |||
| Yi-9B | 2024-03-01 | Unverified | yi-9b | ||||||||||||||
| Falcon 2 11B | falcon-11b | Falcon 2-11B | Technology Innovation Institute | United Arab Emirates | 2024-05-09 | Open weights (restricted use) | 2024-05-09 | 3.6e+23 | Confident | Falcon 2 11B | falcon-2-11b | 109.09 | 106.26 | 115.83 | |||
| GPT-3.5 Turbo | gpt-3.5-turbo-0613 | GPT-3.5 Turbo (June 2023) | OpenAI | United States of America | 2023-06-13 | API access | 2023-06-13 | Likely | GPT-3.5 Turbo (June 2023) | GPT-3.5 Turbo (Jun 2023) | GPT-3.5 Turbo (Jun 2023) | gpt-3-5-turbo-jun-2023 | 112.90 | 108.79 | 121.52 | ||
| GPT-4 (Mar 2023) | gpt-4-32k-0314 | OpenAI | United States of America | 2023-03-14 | API access | 2023-03-15 | 2.1e+25 | Likely | GPT-4 (Mar 2023) | gpt-4-mar-2023 | 125.96 | 121.50 | 133.96 | ||||
| Nemotron-4 15B | Nemotron-4 15B | NVIDIA | United States of America | 2024-02-26 | Unreleased | 2024-02-27 | 7.5e+23 | Confident | A 15 billion-parameter language model trained by Nvidia. | Nemotron-4 15B | nemotron-4-15b | 106.82 | 103.05 | 111.71 | |||
| Megatron-Turing NLG 530B | Megatron-Turing NLG 530B | Microsoft,NVIDIA | United States of America | 2022-01-28 | Unreleased | 2021-10-11 | 8.6e+23 | Confident | Megatron-Turing NLG 530B | megatron-turing-nlg-530b | |||||||
| INTELLECT-1 | INTELLECT-1-Instruct | Prime Intellect,Hugging Face,Arcee AI | United States of America | 2024-11-29 | 2024-11-29 | 6.0e+22 | Confident | INTELLECT-1 | intellect-1 | 100.43 | 95.53 | 103.62 | |||||
| BLIP-2 (Q-Former) | blip2-opt-2.7b | Salesforce Research | United States of America | 2023-02-06 | Open weights (unrestricted) | 2023-01-30 | 1.2e+21 | Confident | BLIP-2 (Q-Former) | blip-2-q-former | |||||||
| Phi-3.5-vision-instruct | 2024-08-16 | Unverified | phi-3-5-vision-instruct | ||||||||||||||
| MM1-3B-Chat | 2024-03-14 | Unverified | mm1-3b-chat | ||||||||||||||
| MM1-7B-Chat | 2024-03-14 | Unverified | mm1-7b-chat | ||||||||||||||
| llava-v1.6-vicuna-7b | 2024-01-31 | Unverified | llava-v1-6-vicuna-7b | ||||||||||||||
| llama3-llava-next-8b | 2024-04-20 | Unverified | llama3-llava-next-8b | ||||||||||||||
| Gemini 1.0 Pro Vision | gemini-1.0-pro-vision | Google DeepMind | United States of America | 2024-01-04 | API access | 2024-02-15 | Unknown | A vision language model from Google DeepMind. | Gemini 1.0 Pro Vision | gemini-1-0-pro-vision | |||||||
| falcon-11B-vlm | 2024-05-21 | Unverified | falcon-11b-vlm | ||||||||||||||
| llava-v1.6-vicuna-13b | 2024-01-31 | Unverified | llava-v1-6-vicuna-13b | ||||||||||||||
| llava-v1.6-mistral-7b | 2024-01-31 | Unverified | llava-v1-6-mistral-7b | ||||||||||||||
| instructblip-vicuna-7b | 2023-05-22 | Unverified | instructblip-vicuna-7b | ||||||||||||||
| instructblip-vicuna-13b | 2023-12-25 | Unverified | instructblip-vicuna-13b | ||||||||||||||
| InternVL-Chat-ViT-6B-Vicuna-7B | 2023-12-25 | Unverified | internvl-chat-vit-6b-vicuna-7b | ||||||||||||||
| InternVL-Chat-ViT-6B-Vicuna-13B | 2024-08-16 | Unverified | internvl-chat-vit-6b-vicuna-13b | ||||||||||||||
| llava-v1.5-7b | 2023-10-05 | Unverified | llava-v1-5-7b | ||||||||||||||
| DeepSeek-R1 (May 2025) | chutes/DeepSeek-R1-0528 | DeepSeek-R1 (May 2025) | DeepSeek | China | 2025-05-28 | Open weights (unrestricted) | 2025-05-28 | 4.0e+24 | Confident | DeepSeek-R1 (May 2025) (Chutes) | DeepSeek-R1 (May 2025) | deepseek-r1-may-2025 | 142.19 | 139.19 | 148.16 | ||
| chutes/GLM-4.5-FP8 | 2025-07-27 | Unverified | chutes-glm-4-5-fp8 | ||||||||||||||
| gpt-oss-120b | chutes/gpt-oss-120b | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | gpt-oss-120b (Chutes) | gpt-oss-120b | gpt-oss-120b | 140.79 | 135.32 | 148.00 | |||
| gpt-oss-120b | chutes/gpt-oss-120b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | gpt-oss-120b (high) (Chutes) | gpt-oss-120b | gpt-oss-120b | 140.79 | 135.32 | 148.00 | |||
| Llama 4 Maverick | chutes/Llama-4-Maverick-17B-128E-Instruct | Meta AI | United States of America | 2025-04-06 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | Likely | Llama 4 Maverick (128 experts) (Chutes) | Llama 4 Maverick | llama-4-maverick | 133.13 | 129.73 | 138.87 | |||
| Qwen3-8B | chutes/Qwen3-8B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 1.8e+24 | Confident | Qwen3 8B (Chutes) | Qwen3-8B | qwen3-8b | ||||||
| Qwen3-235B-A22B | chutes/Qwen3-235B-A22B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Likely | Qwen3 (Apr 2025) (Chutes) | Qwen3-235B-A22B | qwen3-235b-a22b | 139.73 | 135.99 | 145.21 | |||
| Qwen3-235B-A22B (Jul 2025) | chutes/Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | Qwen3 (Jul 2025) (Chutes) | Qwen3-235B-A22B-Thinking (Jul 2025) | Qwen3-235B-A22B-Thinking (Jul 2025) | qwen3-235b-a22b-thinking-jul-2025 | 145.81 | 142.32 | 152.99 | ||
| QwQ-32B | chutes/QwQ-32B | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B (Chutes) | QwQ-32B | QwQ-32B | qwq-32b | |||||
| deepinfra/Qwen3-Next-80B-A3B-Instruct | 2025-09-11 | Unverified | deepinfra-qwen3-next-80b-a3b-instruct | ||||||||||||||
| DeepSeek-V3.2-Exp | DeepSeek-V3.2-Exp_high | DeepSeek | China | 2025-09-29 | Open weights (unrestricted) | 2025-09-29 | 4.2e+24 | Confident | DeepSeek-V3.2-Exp | deepseek-v3-2-exp | 145.19 | 141.83 | 151.64 | ||||
| Gemini 2.5 Flash-Lite | gemini-2.5-flash-lite-preview-06-17-thinking | gemini-2.5-flash-lite-preview-06-17 (Default thinking length) | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-15 | Unknown | A June 2025 preview of Gemini 2.5 Flash Lite. | Gemini 2.5 Flash-Lite (Jun 2025) | Gemini 2.5 Flash-Lite (Jun 2025) | Gemini 2.5 Flash-Lite (Jun 2025) | gemini-2-5-flash-lite-jun-2025 | ||||
| DeepSeek-V3 (Mar 2025) | chutes/DeepSeek-V3-0324 | DeepSeek-V3 (Mar 2025) | DeepSeek | China | 2025-03-24 | Open weights (restricted use) | 2025-03-24 | 3.3e+24 | Confident | DeepSeek-V3 (Mar 2025) (Chutes) | DeepSeek-V3 (Mar 2025) | deepseek-v3-mar-2025 | 137.19 | 134.71 | 142.64 | ||
| Gemma 3 27B | chutes/Gemma-3-27b-It | Google DeepMind | United States of America | 2025-03-11 | Open weights (restricted use) | 2025-03-12 | 2.3e+24 | Confident | Gemma-3-27b-it (Chutes) | Gemma 3 27B | gemma-3-27b | 131.14 | 127.45 | 136.64 | |||
| Kimi K2 | fireworks/Kimi-K2-Instruct-0905 | Kimi K2 Instruct (0905) | Moonshot | China | 2025-09-05 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | Kimi K2 (Sep 2025) | Kimi K2 (Sep 2025) | kimi-k2-sep-2025 | 141.16 | 138.25 | 147.42 | ||
| Grok-3 mini | grok-3-mini-beta_medium | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | Unknown | A beta release of a smaller version of XAI's third generation Grok model, evaluated with medium reasoning effort. | Grok-3 mini | grok-3-mini | 141.19 | 138.81 | 147.60 | ||||
| MiniMax-M1-80k | MiniMax-M1-80k | MiniMax | China | 2025-06-13 | Open weights (unrestricted) | 2025-06-13 | 4.3e+24 | Likely | MiniMax-M1-80k | minimax-m1-80k | |||||||
| nvidia-nemotron-nano-9b-v2 | 2025-08-18 | Unverified | nvidia-nemotron-nano-9b-v2 | ||||||||||||||
| Qwen3-235B-A22B (Jul 2025) | parasail-qwen3-235b-a22b-instruct-2507 | Alibaba | China | 2025-09-01 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | Qwen3 (Parasail) | Qwen3-235B-A22B-Instruct (Jul 2025) | Qwen3-235B-A22B-Instruct (Jul 2025) | qwen3-235b-a22b-instruct-jul-2025 | 139.11 | 135.96 | 144.57 | ||
| Qwen3-30B-A3B | chutes/Qwen3-30B-A3B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 6.5e+23 | Likely | Qwen3-30B-A3B (Chutes) | Qwen3-30B-A3B | qwen3-30b-a3b | ||||||
| Qwen3-14B | chutes/Qwen3-14B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 3.2e+24 | Confident | Qwen3 14B | Qwen3-14B | qwen3-14b | ||||||
| Qwen3-32B | chutes/Qwen3-32B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 7.1e+24 | Confident | Qwen3 32B (Chutes) | Qwen3-32B | qwen3-32b | ||||||
| chutes/Qwen3-Next-80B-A3B-Instruct | 2025-09-11 | Unverified | chutes-qwen3-next-80b-a3b-instruct | ||||||||||||||
| Llama 4 Scout | chutes/Llama-4-Scout-17B-16E Instruct | Meta AI | United States of America | 2025-04-06 | Open weights (restricted use) | 2025-04-05 | 4.1e+24 | Likely | Llama 4 Maverick (16 experts) (Chutes) | Llama 4 Scout | llama-4-scout | 130.61 | 127.27 | 135.54 | |||
| GPT-3 Medium | ada | OpenAI | United States of America | 2020-06-22 | 6.4e+20 | Confident | GPT-3 Medium | gpt-3-medium | |||||||||
| GPT-3 XL | babbage | OpenAI | United States of America | 2020-06-22 | 2020-06-22 | 2.4e+21 | Confident | GPT-3 XL (babbage) | GPT-3 XL (babbage) | gpt-3-xl-babbage | |||||||
| Baichuan 2-7B | Baichuan-2-7B-Base | Baichuan | China | 2023-09-20 | Open weights (restricted use) | 2023-09-20 | 1.1e+23 | Confident | Baichuan 2-7B | baichuan-2-7b | 95.63 | 89.73 | 101.68 | ||||
| Baichuan2-13B | Baichuan-2-13B-Base | Baichuan | China | 2023-09-06 | Open weights (restricted use) | 2023-09-06 | 2.0e+23 | Confident | Baichuan2-13B | baichuan2-13b | 102.40 | 93.75 | 108.78 | ||||
| BLOOM-176B | bloom | Hugging Face,BigScience | France,United States of America | 2022-07-06 | Open weights (restricted use) | 2022-07-11 | 3.7e+23 | Confident | BLOOM-176B | bloom-176b | |||||||
| chatglm2-6b | 2023-06-24 | Unverified | chatglm2-6b | ||||||||||||||
| GPT-3 6.7B | curie | OpenAI | United States of America | 2020-06-22 | 1.2e+22 | Confident | GPT-3 6.7B | gpt-3-6-7b | |||||||||
| GPT-3 175B (davinci) | davinci | OpenAI | United States of America | API access | 2020-05-28 | 3.1e+23 | Confident | GPT-3 175B (davinci) | GPT-3 175B (davinci) | gpt-3-175b-davinci | |||||||
| GPT-3.5 (davinci-003) | text-davinci-003 | 2022-11-28 | 2022-11-28 | Likely | davinci-003 | davinci-003 | davinci-003 | ||||||||||
| internlm-7b | 2023-07-05 | Unverified | internlm-7b | ||||||||||||||
| internlm-20b | 2023-09-18 | Unverified | internlm-20b | ||||||||||||||
| OPT-66B | opt-66b | Meta AI | United States of America | 2022-05-03 | Open weights (non-commercial) | 2022-06-21 | 1.1e+23 | Confident | Facebook's 66 billion-parameter model in the OPT series. | OPT-66B | opt-66b | ||||||
| OPT-175B | opt-175b | Meta AI | United States of America | 2022-05-02 | Open weights (non-commercial) | 2022-05-02 | 4.3e+23 | Confident | Facebook's 66 billion-parameter model in the OPT series. | OPT-175B | opt-175b | ||||||
| Stable Beluga 2 | StableBeluga2 | Stability AI | United Kingdom of Great Britain and Northern Ireland | 2023-07-20 | Open weights (non-commercial) | 2023-07-20 | Likely | Stable Beluga 2 | Stable Beluga 2 | stable-beluga-2 | 116.35 | 113.74 | 119.46 | ||||
| InstructGPT 350M | text-ada-001 | OpenAI | United States of America | 2022-01-27 | Confident | InstructGPT 350M | instructgpt-350m | ||||||||||
| InstructGPT 1.3B | text-babbage-001 | OpenAI | United States of America | API access | 2022-01-27 | Confident | InstructGPT 1.3B | instructgpt-1-3b | |||||||||
| InstructGPT 6B | text-curie-001 | OpenAI | United States of America | API access | 2022-01-27 | Confident | InstructGPT 6B | instructgpt-6b | |||||||||
| GPT-5.2 Pro | gpt-5.2-pro-2025-12-11_xhigh | GPT-5.2 Pro | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Likely | GPT-5.2 Pro (xhigh) | GPT-5.2 Pro | gpt-5-2-pro | 154.19 | 150.15 | 162.75 | |||
| o1 | o1-2024-12-17_low | o1 (low) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | Unknown | OpenAI's first reasoning model, evaluated with at the low reasoning tokens level. | o1 | o1 | 142.83 | 139.83 | 148.65 | |||
| o1-pro | o1-pro-2025-03-19_low | o1 Pro (low) | OpenAI | United States of America | 2025-03-19 | API access | 2025-03-19 | Likely | An An enhanced version of OpenAI's first reasoning model,o1, evaluated with low reasoning effort. | o1-pro | o1-pro | ||||||
| Claude Opus 4.6 | claude-opus-4-6_high | Claude Opus 4.6 (high) | Anthropic | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | Claude Opus 4.6 (high) | Claude Opus 4.6 | claude-opus-4-6 | 155.28 | 152.40 | 163.08 | |||
| Reka Flash 3 | reka-flash-3 | Reka AI | United States of America | 2025-03-10 | Open weights (unrestricted) | 2025-03-10 | Confident | Reka Flash 3 | Reka Flash 3 | reka-flash-3 | |||||||
| Mistral NeMo | Mistral-Nemo-Instruct-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | Confident | Mistral NeMo | mistral-nemo | 118.27 | 114.37 | 123.07 | |||||
| Llama 3.2 11B | Llama-3.2-11B-Vision-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 5.8e+23 | Confident | An 11 billion-parameter instruction-tuned vision language model in Meta's Llama 3.2 series. | Llama 3.2 11B | llama-3-2-11b | ||||||
| Llama 3.2 3B | Llama-3.2-3B-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 1.7e+23 | Confident | A 3 billion-parameter instruction-tuned model in Meta's Llama 3.2 series. | Llama 3.2 3B | llama-3-2-3b | ||||||
| Qwen2.5-7B | qwen2.5-7b-instruct | Alibaba | China | 2025-02-26 | Open weights (unrestricted) | 2024-09-19 | 8.2e+23 | Confident | An instruction-tuned version of Qwen 2.5 7b. | Qwen2.5-7B | Qwen2.5-7B | qwen2-5-7b | |||||
| Llama 3.2 1B | Llama-3.2-1B-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 6.6e+22 | Confident | A 1 billion-parameter instruction-tuned model in Meta's Llama 3.2 series. | Llama 3.2 1B | llama-3-2-1b | ||||||
| GPT-5.3 Codex | gpt-5.3-codex_xhigh | GPT-5.3 Codex (xhigh) | OpenAI | United States of America | 2026-02-05 | API access | 2026-02-05 | Likely | GPT-5.3 Codex | gpt-5-3-codex | 155.89 | 153.05 | 164.24 | ||||
| GPT-5.5 | gpt-5.5_none | GPT-5.5 (no thinking) | OpenAI | United States of America | 2026-04-23 | API access | 2026-04-23 | Likely | GPT-5.5 | gpt-5-5 | 158.22 | 155.43 | 165.84 | ||||
| Claude Sonnet 4.6 | claude-sonnet-4-6_medium | Claude Sonnet 4.6 (medium) | Anthropic | United States of America | 2026-02-17 | API access | 2026-02-17 | Likely | Claude Sonnet 4.6 | claude-sonnet-4-6 | 152.58 | 148.95 | 158.53 | ||||
| GPT-5.4 | gpt-5.4-2026-03-05_none | GPT-5.4 (none) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 (none) | GPT-5.4 | gpt-5-4 | 156.13 | 153.44 | 163.96 | |||
| GPT-5.4 Pro | gpt-5.4-pro-2026-03-05_none | GPT-5.4 Pro (no thinking) | OpenAI | United States of America | 2026-03-05 | API access | 2026-03-05 | Likely | GPT-5.4 Pro (no thinking) | GPT-5.4 Pro | gpt-5-4-pro | 157.73 | 155.08 | 165.39 | |||
| GPT‑5-Codex | gpt-5-codex_high | GPT-5-codex (high) | OpenAI | United States of America | 2025-09-15 | API access | 2025-09-15 | Unknown | GPT-5-codex | GPT‑5-Codex | gpt5-codex | ||||||
| GPT-5.2 | gpt-5.2-2025-12-11_none | GPT-5.2 (none) | OpenAI | United States of America | 2025-12-11 | API access | 2025-12-11 | Unknown | GPT-5.2 (none) | GPT-5.2 | gpt-5-2 | 153.72 | 150.93 | 161.15 | |||
| DeepSeek-V4-Pro | deepseek-v4-pro_max | DeepSeek v4 Pro (max) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 9.7e+24 | Likely | DeepSeek v4 (max) | DeepSeek-V4-Pro | deepseek-v4-pro | |||||
| DeepSeek-V4-Pro | deepseek-v4-pro_high | DeepSeek v4 Pro (high) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 9.7e+24 | Likely | DeepSeek v4 (high) | DeepSeek-V4-Pro | deepseek-v4-pro | |||||
| DeepSeek-V4-Flash | deepseek-v4-flash_max | DeepSeek v4 Flash (max) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 2.5e+24 | Likely | DeepSeek v4 (max) | DeepSeek-V4-Flash | deepseek-v4-flash | |||||
| DeepSeek-V4-Flash | deepseek-v4-flash_high | DeepSeek v4 Flash (high) | DeepSeek | China | 2026-04-24 | Open weights (unrestricted) | 2026-04-24 | 2.5e+24 | Likely | DeepSeek v4 (max) | DeepSeek-V4-Flash | deepseek-v4-flash | |||||
| Grok-3 mini | grok-3-mini_high | Grok 3 mini | xAI | United States of America | 2025-06-24 | API access | 2025-02-19 | Unknown | Grok 3 mini (high) | Grok-3 mini | grok-3-mini | 141.19 | 138.81 | 147.60 | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-09-2025_16k | Google DeepMind | United States of America | 2025-09-25 | API access | 2025-04-17 | Unknown | Gemini 2.5 Flash-Lite (Sep 2025, 16k thinking) | Gemini 2.5 Flash (Sep 2025) | Gemini 2.5 Flash (Sep 2025) | gemini-2-5-flash-sep-2025 | 143.39 | 139.01 | 151.42 | |||
| gpt-oss-120b | gpt-oss-120b_medium | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | The larger one of OpenAI's two latest open source language models, evaluated with medium reasoning effort. | gpt-oss-120b | gpt-oss-120b | 140.79 | 135.32 | 148.00 | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-04-17 (16K thinking) | Google DeepMind | United States of America | 2025-04-17 | API access | 2025-04-17 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 16,000 reasoning tokens. | Gemini 2.5 Flash (Apr 2025) | Gemini 2.5 Flash (Apr 2025) | gemini-2-5-flash-apr-2025 | 140.71 | 137.20 | 146.83 | |||
| gpt-oss-20b | gpt-oss-20b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | Confident | The smaller one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-20b | gpt-oss-20b | ||||||
| GLM-4.5 | glm-4.5_thinking | GLM-4.5 Thinking | Z.ai (Zhipu AI),Tsinghua University | China | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.4e+24 | Confident | GLM-4.5 | glm-4-5 | ||||||
| GPT-5 | gpt-5-chat | OpenAI | United States of America | API access | 2025-08-07 | 6.6e+25 | Speculative | GPT-5 | gpt-5 | 150.00 | 146.38 | 157.62 | |||||
| Kimi K2 | kimi-k2-0711-preview | Moonshot | China | 2025-07-12 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | A preview of an updated version of Kimi K2. | Kimi K2 (Jul 2025) | Kimi K2 (Jul 2025) | kimi-k2-jul-2025 | 140.60 | 137.19 | 147.28 | ||
| Qwen3-235B-A22B (Jul 2025) | qwen/qwen3-235b-a22b-thinking-2507 | Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | Qwen3-235B-A22B-Thinking (Jul 2025) | Qwen3-235B-A22B-Thinking (Jul 2025) | qwen3-235b-a22b-thinking-jul-2025 | 145.81 | 142.32 | 152.99 | ||
| GPT-5.4 Nano | gpt-5.4-nano-2026-03-17_none | GPT-5.4 nano (no thinking) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 nano (high) | GPT-5.4 Nano | gpt-5-4-nano | 146.23 | 142.65 | 153.80 | |||
| GPT-5.4 Mini | gpt-5.4-mini-2026-03-17_none | GPT-5.4 mini (none) | OpenAI | United States of America | 2026-03-17 | API access | 2026-03-17 | Likely | GPT-5.4 mini (high) | GPT-5.4 Mini | gpt-5-4-mini | 148.74 | 145.69 | 155.92 | |||
| gpt-oss-20b | gpt-oss-20b_medium | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | Confident | The smaller one of OpenAI's two latest open source language models, evaluated with medium reasoning effort. | gpt-oss-20b | gpt-oss-20b | ||||||
| Kimi K2 | Kimi-K2-Instruct-0905 | Moonshot | China | 2025-09-05 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | Kimi K2 (Sep 2025) | Kimi K2 (Sep 2025) | kimi-k2-sep-2025 | 141.16 | 138.25 | 147.42 | |||
| Gemini 2.5 Flash-Lite | gemini-2.5-flash-lite-preview-06-17_16K | Google DeepMind | United States of America | 2025-06-17 | API access | 2025-06-15 | Unknown | A June 2025 preview of Gemini 2.5 Flash Lite, evaluated with up to 16,000 reasoning tokens. | Gemini 2.5 Flash-Lite (Jun 2025, 16k thinking) | Gemini 2.5 Flash-Lite (Jun 2025) | Gemini 2.5 Flash-Lite (Jun 2025) | gemini-2-5-flash-lite-jun-2025 | |||||
| qwen3-coder-next | 2026-02-02 | Unverified | qwen3-coder-next | ||||||||||||||
| mistral-medium-2508 | 2025-08-12 | Unverified | mistral-medium-2508 | ||||||||||||||
| Qwen-1_8B | Qwen-1.8B | 2023-11-30 | Unverified | qwen-1-8b | |||||||||||||
| Qwen-7B | Qwen-7B | Alibaba | China | 2023-09-28 | Open weights (restricted use) | 2023-09-28 | 1.0e+23 | Confident | Qwen's 7 billion-parameter model in the original Qwen series. | Qwen-7B | Qwen-7B | qwen-7b | 106.46 | 99.77 | 111.36 | ||
| Qwen-14B | Qwen-14B | Alibaba | China | 2023-09-28 | Open weights (restricted use) | 2023-09-28 | 2.5e+23 | Confident | Qwen's 14 billion-parameter model in the original Qwen series. | Qwen-14B | Qwen-14B | qwen-14b | 112.37 | 108.75 | 116.68 | ||
| Yi-34B | Yi-34B | 01.AI | China | 2023-11-02 | Open weights (restricted use) | 2023-11-02 | 6.1e+23 | Confident | Yi-34B | Yi-34B | Yi-34B | yi-34b | 117.40 | 114.48 | 121.38 | ||
| Baichuan1-7B | Baichuan-7B | Baichuan | China | 2023-06-01 | Open weights (non-commercial) | 2023-06-01 | 5.0e+22 | Confident | Baichuan1-7B | baichuan1-7b | 89.81 | 85.07 | 95.64 | ||||
| Baichuan 1-13B | Baichuan-13B-Base | Baichuan | China | 2023-07-11 | Open weights (restricted use) | 2023-07-11 | 9.4e+22 | Confident | Baichuan 1-13B | baichuan-1-13b | |||||||
| Llama 2-13B | Llama-2-13b-chat | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 1.6e+23 | Confident | A chat-optimized version of Meta's 13 billion-parameter model in the Llama 2 series. | Llama 2-13B | llama-2-13b | 105.26 | 100.89 | 108.96 | |||
| Llama 2-70B | Llama-2-70b-chat | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | Confident | A chat-optimized version of Meta's 70 billion-parameter model in the Llama 2 series. | Llama 2-70B | llama-2-70b | 113.14 | 109.14 | 116.76 | |||
| Baichuan2-13B-Chat | 2023-09-06 | Unverified | baichuan2-13b-chat | ||||||||||||||
| Qwen-14B | Qwen-14B-Chat | Alibaba | China | 2023-09-24 | Open weights (restricted use) | 2023-09-28 | 2.5e+23 | Confident | A chat-optimized version of Qwen's first-generation 14 billion-parameter model. | Qwen-14B (chat) | Qwen-14B | qwen-14b | 112.37 | 108.75 | 116.68 | ||
| internlm-chat-20b | 2023-09-17 | Unverified | internlm-chat-20b | ||||||||||||||
| Yi 6B | Yi-6B-Chat | 01.AI | China | 2023-11-22 | Open weights (restricted use) | 2023-11-02 | 1.3e+23 | Confident | Yi-6B | Yi 6B | yi-6b | 104.06 | 98.93 | 108.63 | |||
| Inflection-1 | Inflection-1 | Inflection AI | United States of America | 2023-06-22 | Hosted access (no API) | 2023-06-23 | 1.0e+24 | Speculative | Inflection-1 | inflection-1 | |||||||
| Gemini 1.5 Flash | gemini-1.5-flash-0514 | Google DeepMind | United States of America | 2024-05-14 | API access | 2024-05-10 | Unknown | Gemini 1.5 Flash (May 2024) | Gemini 1.5 Flash (May 2024) | gemini-1-5-flash-may-2024 | 122.52 | 118.82 | 126.23 | ||||
| Mistral 7B | Mistral-7B-Instruct-v0.2 | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-10-10 | Confident | Mistral 7B v0.2 | Mistral 7B v0.2 | mistral-7b-v0-2 | |||||||
| mistral-small-2402 | 2024-02-26 | Unverified | mistral-small-2402 | ||||||||||||||
| Mixtral 8x22B | Mixtral-8x22B-v0.1 | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | Speculative | Mixtral 8x22B | mixtral-8x22b | 121.21 | 114.22 | 125.22 | ||||
| Qwen1.5-7B | Qwen1.5-7B | Alibaba | China | 2024-02-04 | Open weights (unrestricted) | 2024-02-04 | 1.7e+23 | Confident | A 7 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-7B | Qwen1.5-7B | qwen1-5-7b | |||||
| Qwen1.5-14B | qwen1.5-14B | Alibaba | China | 2024-02-04 | Open weights (unrestricted) | 2024-02-04 | 3.4e+23 | Confident | A 14 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-14B | Qwen1.5-14B | qwen1-5-14b | |||||
| Qwen1.5-32B | qwen1.5-32B | Alibaba | China | 2024-02-04 | Open weights (restricted use) | 2024-02-05 | Confident | A 32 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-32B | Qwen1.5-32B | qwen1-5-32b | ||||||
| Qwen2.5-14B | qwen2.5-14b-instruct | Alibaba | China | 2025-02-26 | Open weights (unrestricted) | 2024-09-19 | 1.6e+24 | Confident | An instruction-tuned version of Qwen 2.5 14b. | Qwen2.5-14B (instruct) | Qwen2.5-14B | qwen2-5-14b | |||||
| Yi-Large | Yi-large | 01.AI | China | 2024-05-13 | API access | 2024-05-13 | 1.8e+24 | Speculative | Yi-Large | Yi-Large | yi-large | ||||||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_15K | Claude 3.7 Sonnet (15k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 15,000 reasoning tokens allowed. | Claude 3.7 Sonnet (15k thinking) | Claude 3.7 Sonnet | claude-3-7-sonnet | 142.03 | 139.52 | 148.45 | |
| Pixtral 12B | Pixtral-12B-2409 | Mistral AI | France | 2024-09-17 | Open weights (unrestricted) | 2024-09-17 | Confident | Pixtral 12B | pixtral-12b | ||||||||
| ml-elephant | Zoo Ml-Ephant | Unverified | ml-elephant | ||||||||||||||
| MPT-30B | mpt-30b-instruct | MosaicML | United States of America | 2023-06-22 | Open weights (unrestricted) | 2023-06-22 | 1.9e+23 | Confident | MPT-30B | mpt-30b | 99.88 | 96.89 | 103.05 | ||||
| Vicuna-13B-v1.3 | vicuna-13b-v1.3 | Large Model Systems Organization,University of California (UC) Berkeley | United States of America | 2023-06-18 | Open weights (restricted use) | 2023-06-22 | Confident | Vicuna-13B-v1.3 | Vicuna-13B-v1.3 | vicuna-13b-v1-3 | |||||||
| phi-3.5-mini | Phi-3.5-mini-instruct | Microsoft | United States of America | 2024-08-16 | Open weights (unrestricted) | 2024-04-23 | 3.7e+22 | Confident | An instruction-tuned version of phi-3.5-mini, a small model in Microsoft's Phi 3.5 open source model series. | phi-3.5-mini | phi-3-5-mini | ||||||
| Phi-3.5-MoE | Phi-3.5-MoE-instruct | Microsoft | United States of America | 2024-08-17 | Open weights (unrestricted) | 2024-04-23 | 3.0e+23 | Confident | Phi-3.5-MoE | phi-3-5-moe | |||||||
| Mistral NeMo | Mistral-Nemo-Base-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | Confident | Mistral NeMo | mistral-nemo | 118.27 | 114.37 | 123.07 | |||||
| Gemma 2 9B | gemma-2-9b | Google DeepMind | United States of America | Open weights (restricted use) | 2024-06-24 | 4.3e+23 | Confident | A 9 billion-parameter open source model from Google DeepMind in the Gemma 2 series. | Gemma 2 9B | gemma-2-9b | 119.51 | 116.30 | 122.43 | ||||
| DeepSeek-V3 | DeepSeek-V3-Base | DeepSeek | China | 2024-12-26 | Open weights (restricted use) | 2024-12-24 | 3.3e+24 | Confident | DeepSeek-V3 (base) | DeepSeek-V3 | deepseek-v3 | 133.13 | 130.21 | 137.90 | |||
| Mistral 7B | Mistral-7B-Instruct-v0.1 | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | Confident | Mistral 7B v0.1 | Mistral 7B v0.1 | mistral-7b-v0-1 | 111.36 | 108.25 | 114.82 | ||||
| GPT-3.5 (davinci-002) | code-davinci-002 | OpenAI | United States of America | 2022-03-15 | API access | 2022-03-15 | 2.6e+24 | Speculative | davinci-002 | davinci-002 | davinci-002 | ||||||
| Falcon-40B | falcon-40b-instruct | Technology Innovation Institute | United Arab Emirates | 2023-05-25 | Open weights (unrestricted) | 2023-03-15 | 2.4e+23 | Confident | Falcon-40B | falcon-40b | 103.53 | 100.09 | 108.57 | ||||
| DeepSeek-Coder-V2-Lite-Instruct | 2024-06-13 | Unverified | deepseek-coder-v2-lite-instruct | ||||||||||||||
| DeepSeek-Coder-V2 236B | DeepSeek-Coder-V2-Instruct | DeepSeek | China | 2024-06-17 | Open weights (restricted use) | 2024-06-17 | 1.3e+24 | Confident | An instruction-tuned version of DeepSeek Coder V2, a 236 billion-parameter coding-optimized model from DeepSeek. | DeepSeek-Coder-V2 (instruct) | DeepSeek-Coder-V2 236B | deepseek-coder-v2-236b | |||||
| Qwen2.5-Coder-3B-Instruct | 2024-11-06 | Unverified | qwen2-5-coder-3b-instruct | ||||||||||||||
| Qwen2.5-Coder (7B) | Qwen2.5-Coder-7B-Instruct | Alibaba | China | 2024-09-17 | Open weights (unrestricted) | 2024-09-18 | 2.5e+23 | Confident | An instruction-tuned and coding-optimized 7 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-Coder-7B (instruct) | Qwen2.5-Coder (7B) | qwen2-5-coder-7b | 112.54 | 104.01 | 118.82 | ||
| Qwen2.5-Coder-14B-Instruct | 2024-11-06 | Unverified | qwen2-5-coder-14b-instruct |
This gap between open and closed models is slightly larger than the one we identified in our October 2025 Data Insight, which found that open models lagged by an average of three months between January 2023 and October 2025.
Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.
Learn more about this graph
We calculate the average amount of time it takes for the best open-weight models to catch up with state-of-the-art performance according to our internal capability metric, the Epoch Capability Index (ECI). ECI is a composite measure that captures performance across many benchmarks.
Analysis
To calculate the average time gap, we proceed day by day over our analysis window, January 1, 2026, through May 28, 2026. For each day, we identify the open-weight model with the highest ECI score available by that date. We then compare that model to the historical state-of-the-art ECI frontier and ask: what is the most recent date on which the SOTA model was not significantly better than this open-weight model? The time gap for that day is the number of days elapsed since that date.
Because ECI scores are estimated with uncertainty, we use bootstrap samples to make this comparison. Bootstrap samples are generated by resampling our full set of benchmark scores with replacement and refitting the ECI model on each resampled dataset. For each bootstrap sample, we compare the open-weight model’s bootstrapped ECI estimate to the bootstrapped ECI estimate of each historical SOTA model, preserving the pairing between bootstrap samples across models. We treat the open-weight model as having plausibly caught up to a previous SOTA model if it outperforms that SOTA model in at least 5% of paired bootstrap samples. Equivalently, this means the previous SOTA model is not significantly better than the open-weight model at the 5% level. We then use the most recent historical SOTA date satisfying this criterion to compute the time gap.
We find an average time gap of four months. This estimate would grow to six months if we required that the open-weight model’s point estimate for ECI be strictly higher than the closed-weight model it is catching up to (instead of better in at least 5% of samples).
Because release dates are observed without uncertainty, calculating the average ECI gap (i.e. the “vertical” gap at a given date) is as simple as observing the average difference between absolute and open-weight SOTA across all dates in our analysis window. We find an average ECI gap of 8 points, with a 90% confidence interval of 7 to 11 units.
Limitations
Two factors mean that our estimate may tend to understate the true gap between open and closed models.
First, evidence suggests that open-weight models tend to perform worse on private benchmarks compared to closed models, plausibly because they more aggressively hillclimb on public benchmarks.
Second, we only include models with enough public benchmark coverage to assign an ECI. Leading closed labs do not always release their most capable models, for safety, commercial, or competitive reasons.

