Since 2023, every model at the frontier of AI capabilities, as measured by the Epoch Capabilities Index, has been developed in the United States. Over that same period, Chinese models have trailed US capabilities by an average of seven months, with a minimum gap of four months and a maximum gap of 14.
| Model | Model version ID | Display name | Organization | Country (of organization) | Version release date | Model accessibility | Publication date | Training compute (FLOP) | Confidence | Description | Unique display name | Slug | ECI | ECI CI low | ECI CI high |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.2 | gpt-5.2-pro-2025-12-11-webapp | GPT-5.2 (Pro) | OpenAI | United States of America | 2025-12-11 | 2025-12-11 | Unverified | GPT-5.2 (Pro) | gpt-5-2 | 152.03 | 148.77 | 157.16 | |||
| DeepSeek-V3.2 | fireworks/deepseek-v3p2 | DeepSeek | China | 2025-12-01 | 2025-12-01 | Unverified | deepseek-v3-2 | 144.60 | 140.97 | 147.18 | |||||
| Gemini 2.5 Flash (Jun 2025) | gemini-2.5-flash | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-06-17 | API access | 2025-06-17 | Unknown | A small model from Google DeepMind's Gemini 2.5 series. | Gemini 2.5 Flash (Jun 2025) | gemini-2-5-flash-jun-2025 | |||||
| Gemini 3 Flash | gemini-3-flash-preview | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-12-17 | 2025-12-17 | Unverified | gemini-3-flash | 152.87 | 150.33 | 156.36 | |||||
| DeepSeek-V3.2 | deepseek-reasoner | DeepSeek | China | 2025-12-01 | 2025-12-01 | Unverified | deepseek-v3-2 | 144.60 | 140.97 | 147.18 | |||||
| gpt-oss-120b | openai/gpt-oss-120b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | gpt-oss-120b (high) | gpt-oss-120b | 139.64 | 135.24 | 143.10 | ||
| GPT-5.2 | gpt-5.2-2025-12-11_xhigh | GPT-5.2 (xhigh) | OpenAI | United States of America | 2025-12-11 | 2025-12-11 | Unverified | GPT-5.2 (xhigh) | gpt-5-2 | 152.03 | 148.77 | 157.16 | |||
| Qwen3-235B-A22B-Thinking (Jul 2025) | Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | A 22 billion active (235 billion total)-parameter mixture-of-experts reasoning model in the Qwen 3 series. | qwen3-235b-a22b-thinking-jul-2025 | 145.26 | 141.93 | 148.42 | ||
| GPT-5.2 | gpt-5.2-2025-12-11_high | GPT-5.2 (high) | OpenAI | United States of America | 2025-12-11 | 2025-12-11 | Unverified | GPT-5.2 (high) | gpt-5-2 | 152.03 | 148.77 | 157.16 | |||
| GPT-5.2 | gpt-5.2-2025-12-11_medium | GPT-5.2 (medium) | OpenAI | United States of America | 2025-12-11 | 2025-12-11 | Unverified | GPT-5.2 (medium) | gpt-5-2 | 152.03 | 148.77 | 157.16 | |||
| GPT-5.2 | gpt-5.2-2025-12-11_low | GPT-5.2 (low) | OpenAI | United States of America | 2025-12-11 | 2025-12-11 | Unverified | GPT-5.2 (low) | gpt-5-2 | 152.03 | 148.77 | 157.16 | |||
| Qwen3-235B-A22B-Thinking (Jul 2025) | Qwen/Qwen3-235B-A22B-Thinking-2507 | Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | qwen3-235b-a22b-thinking-jul-2025 | 145.26 | 141.93 | 148.42 | ||
| Qwen3-Max | qwen3-max-2025-09-23 | Qwen3-Max-Instruct | Alibaba | China | 2025-09-24 | API access | 2025-09-05 | 1.5e+25 | Speculative | A 1-trillion total parameter scale model in the Qwen 3 series. | Qwen3 Max | qwen3-max | 145.28 | 139.81 | 153.20 |
| Kimi K2 Thinking | kimi-k2-thinking-turbo | kimi-k2-thinking (turbo official) | Moonshot | China | 2025-11-06 | Open weights (restricted use) | 2025-11-06 | 4.2e+24 | Likely | kimi-k2-thinking | 145.05 | 142.80 | 146.99 | ||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_59K | Claude Sonnet 4.5 (59k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (59k thinking) | claude-sonnet-4-5 | 146.42 | 143.54 | 148.24 | ||
| GPT-5.1 | gpt-5.1-2025-11-13_high | GPT-5.1 (high) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (high) | gpt-5-1 | 150.24 | 148.43 | 151.94 | ||
| Claude Opus 4.1 | claude-opus-4-1-20250805_27K | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model, benchmarked with up to 27,000 reasoning tokens. | Claude Opus 4.1 (27k thinking) | claude-opus-4-1 | 143.84 | 140.66 | 146.39 | ||
| GPT-5 mini | gpt-5-mini-2025-08-07_high | GPT-5 mini (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 mini (high) | gpt-5-mini | 144.19 | 141.46 | 146.35 | |
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_32K | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (32k thinking) | claude-haiku-4-5 | 140.45 | 137.19 | 142.48 | |||
| o4-mini | o4-mini-2025-04-16_high | o4-mini (high) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | Unknown | The latest smaller size o-series reasoning model from OpenAI, evaluated with high reasoning effort. | o4-mini | 144.92 | 142.46 | 147.04 | ||
| Claude Opus 4.5 | claude-opus-4-5-20251101_32K | Claude Opus 4.5 (32k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (32k thinking) | claude-opus-4-5 | 149.93 | 146.69 | 152.48 | ||
| o3 | o3-2025-04-16_high | o3 (high) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | Unknown | OpenAI's second generation o-series reasoning model, evaluated with high reasoning effort. | o3 | 146.76 | 144.48 | 148.80 | ||
| GPT-5 nano | gpt-5-nano-2025-08-07_high | GPT-5 nano (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 nano (high) | gpt-5-nano | 139.77 | 135.31 | 142.12 | |
| Grok-3 mini | grok-3-mini-beta_high | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | Unknown | A beta release of a smaller version of XAI's third generation Grok model, evaluated with high reasoning effort. | grok-3-mini | 140.70 | 137.75 | 142.61 | |||
| GPT-5 | gpt-5-2025-08-07_high | GPT-5 (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 (high) | gpt-5 | 150.00 | ||
| Kimi K2 Thinking | moonshotai/Kimi-K2-Thinking | Moonshot | China | 2025-11-06 | Open weights (restricted use) | 2025-11-06 | 4.2e+24 | Likely | kimi-k2-thinking | 145.05 | 142.80 | 146.99 | |||
| GLM-4.6 | zai-org/GLM-4.6 | Z.ai (Zhipu AI),Tsinghua University | China | 2025-09-30 | Open weights (unrestricted) | 2025-09-30 | 4.4e+24 | Likely | glm-4-6 | ||||||
| Grok 4 | grok-4-0709 | xAI | United States of America | 2025-07-09 | API access | 2025-07-09 | 5.0e+26 | Speculative | Grok 4 | grok-4 | 147.42 | 145.05 | 149.43 | ||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929 | Claude Sonnet 4.5 (no thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Anthropic's latest mid-size model, part of the Claude 4.5 series. | Claude Sonnet 4.5 (no thinking) | claude-sonnet-4-5 | 146.42 | 143.54 | 148.24 | |
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-06-17 | API access | 2025-06-05 | Unknown | Google’s general-purpose frontier reasoning model. | Gemini 2.5 Pro | gemini-2-5-pro-jun-2025 | 146.18 | 144.16 | 149.21 | ||
| DeepSeek-R1 (May 2025) | DeepSeek-R1-0528 | DeepSeek-R1 (May 2025) | DeepSeek | China | 2025-05-28 | Open weights (unrestricted) | 2025-05-28 | 4.0e+24 | Confident | An updated May 2025 version of DeepSeek's reasoning model, R1. | DeepSeek-R1 (May 2025) | deepseek-r1-may-2025 | 141.75 | 139.03 | 143.81 |
| Gemini 3 Pro | gemini-3-pro-preview | Gemini 3 Pro Preview | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-11-18 | API access | 2025-11-18 | Unknown | Gemini 3 Pro Preview | gemini-3-pro | 153.99 | 151.50 | 159.37 | ||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_32K | Claude Sonnet 4.5 (32k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Anthropic's latest mid-size model, part of the Claude 4.5 series, evaluated with up to 32,000 reasoning tokens. | Claude Sonnet 4.5 (32k thinking) | claude-sonnet-4-5 | 146.42 | 143.54 | 148.24 | |
| o3-mini | o3-mini-2025-01-31_high | o3-mini (high) | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | Unknown | A smaller version of OpenAI's o3 reasoning model, evaluated with high reasoning effort. | o3-mini | 141.00 | 138.02 | 142.82 | ||
| Claude 3.5 Haiku | claude-3-5-haiku-20241022 | Claude 3.5 Haiku (Oct 2024) | Anthropic | United States of America | 2024-10-22 | API access | 2024-10-22 | Unknown | The smallest model in Anthropic’s Claude 3.5 Series. | Claude 3.5 Haiku (Oct 2024) | claude-3-5-haiku | 127.34 | 118.48 | 131.67 | |
| GPT-5.1 | gpt-5.1-2025-11-13_none | GPT-5.1 (no thinking) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (no thinking) | gpt-5-1 | 150.24 | 148.43 | 151.94 | ||
| GPT-5.1 | gpt-5.1-2025-11-13_low | GPT-5.1 (low) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (low) | gpt-5-1 | 150.24 | 148.43 | 151.94 | ||
| Claude Opus 4.5 | claude-opus-4-5-20251101 | Claude Opus 4.5 (no thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (no thinking) | claude-opus-4-5 | 149.93 | 146.69 | 152.48 | ||
| Claude Opus 4.5 | claude-opus-4-5-20251101_16K | Claude Opus 4.5 (16k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (16k thinking) | claude-opus-4-5 | 149.93 | 146.69 | 152.48 | ||
| GPT-5.1 | gpt-5.1-2025-11-13_medium | GPT-5.1 (medium) | OpenAI | United States of America | 2025-11-13 | API access | 2025-11-13 | Unknown | GPT-5.1 (medium) | gpt-5-1 | 150.24 | 148.43 | 151.94 | ||
| o3 | o3-2025-04-16_low | o3 (low) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | Unknown | OpenAI's latest full-size o-series reasoning model release, evaluated on low reasoning effort. | o3 | 146.76 | 144.48 | 148.80 | ||
| o3 | o3-2025-04-16_medium | o3 (medium) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | Unknown | OpenAI's second generation o-series reasoning model, evaluated with medium reasoning effort. | o3 | 146.76 | 144.48 | 148.80 | ||
| o4-mini | o4-mini-2025-04-16_low | o4-mini (low) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | Unknown | The latest smaller size o-series reasoning model from OpenAI, evaluated with low reasoning effort. | o4-mini | 144.92 | 142.46 | 147.04 | ||
| o4-mini | o4-mini-2025-04-16_medium | o4-mini (medium) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | Unknown | The latest smaller size o-series reasoning model from OpenAI, evaluated with medium reasoning effort. | o4-mini | 144.92 | 142.46 | 147.04 | ||
| GPT-5 mini | gpt-5-mini-2025-08-07_medium | GPT-5 mini (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | A smaller version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 mini (medium) | gpt-5-mini | 144.19 | 141.46 | 146.35 | |
| GPT-5 | gpt-5-2025-08-07_medium | GPT-5 (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 (medium) | gpt-5 | 150.00 | ||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_16K | Claude Sonnet 4.5 (16k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Anthropic's latest mid-size model, part of the Claude 4.5 series, evaluated with up to 16,000 reasoning tokens. | Claude Sonnet 4.5 (16k thinking) | claude-sonnet-4-5 | 146.42 | 143.54 | 148.24 | |
| GPT-4 (Jun 2023) | gpt-4-0613 | GPT-4 (Jun 2023) | OpenAI | United States of America | 2023-06-13 | API access | 2023-03-15 | 2.1e+25 | Likely | The June 2023 version of GPT-4, OpenAI's 2023 multimodal flagship model. | gpt-4-jun-2023 | 121.73 | 113.01 | 125.19 | |
| GPT-4 (Mar 2023) | gpt-4-0314 | GPT-4 (Mar 2023) | OpenAI | United States of America | 2023-03-14 | API access | 2023-03-15 | 2.1e+25 | Likely | The March 2023 version of GPT-4, OpenAI's 2023 multimodal flagship model. | gpt-4-mar-2023 | 125.76 | 118.56 | 130.42 | |
| Claude Haiku 4.5 | claude-haiku-4-5-20251001 | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | claude-haiku-4-5 | 140.45 | 137.19 | 142.48 | ||||
| GPT-5 Pro | gpt-5-pro-2025-10-06_high | OpenAI | United States of America | 2025-10-07 | API access | 2025-10-07 | Unknown | gpt-5-pro | 150.34 | 147.37 | 152.54 | ||||
| DeepSeek-V3.1 | DeepSeek-V3.1 | DeepSeek | China | 2025-08-21 | Open weights (unrestricted) | 2025-08-21 | 3.6e+24 | Confident | An updated version of DeepSeek V3. | DeepSeek-V3.1 | deepseek-v3-1 | 138.53 | 135.09 | 142.99 | |
| GLM-4.5 | glm-4.5 | Z.ai (Zhipu AI),Tsinghua University | China | 2025-08-03 | Open weights (unrestricted) | 2025-08-05 | 4.4e+24 | Confident | Zhipu AI & Tsinghua University’s large open source model. | glm-4-5 | |||||
| GPT-5 nano | gpt-5-nano-2025-08-07_medium | GPT-5 nano (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | Unknown | The smallest version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 nano (medium) | gpt-5-nano | 139.77 | 135.31 | 142.12 | |
| GPT-4o (Nov 2024) | gpt-4o-2024-11-20 | GPT-4o (Nov 2024) | OpenAI | United States of America | 2024-11-20 | API access | 2024-11-20 | Speculative | The November 2024 version of GPT-4o, OpenAI's then-flagship multimodal language model. | gpt-4o-nov-2024 | 129.81 | 125.81 | 133.42 | ||
| Claude Opus 4 | claude-opus-4-20250514_27K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 27,000 reasoning tokens. | Claude Opus 4 (27k thinking) | claude-opus-4 | 143.06 | 140.01 | 145.22 | ||
| Claude Opus 4.1 | claude-opus-4-1-20250805 | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model. | Claude Opus 4.1 | claude-opus-4-1 | 143.84 | 140.66 | 146.39 | ||
| Claude Opus 4.1 | claude-opus-4-1-20250805_16K | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Unknown | Anthropic’s largest flagship model, benchmarked with up to 16,000 reasoning tokens. | Claude Opus 4.1 (16k thinking) | claude-opus-4-1 | 143.84 | 140.66 | 146.39 | ||
| GPT-4o mini | gpt-4o-mini-2024-07-18 | OpenAI | United States of America | 2024-07-18 | API access | 2024-07-18 | Speculative | A smaller version of OpenAI's GPT-4o. | gpt-4o-mini | 126.97 | 120.42 | 129.72 | |||
| Claude Opus 4 | claude-opus-4-20250514 | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1). | Claude Opus 4 | claude-opus-4 | 143.06 | 140.01 | 145.22 | |
| Claude Sonnet 4 | claude-sonnet-4-20250514 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized version of Claude 4 series of reasoning models. | Claude Sonnet 4 | claude-sonnet-4 | 142.30 | 139.81 | 144.40 | ||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro-preview-06-05 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-06-05 | API access | 2025-06-05 | Unknown | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model. | Gemini 2.5 Pro Preview (Jun 2025) | gemini-2-5-pro-jun-2025 | 146.18 | 144.16 | 149.21 | |
| GPT-4.1 | gpt-4.1-2025-04-14 | GPT-4.1 | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | Unknown | A coding-optimized GPT-4 series model from OpenAI. | gpt-4-1 | 137.37 | 133.79 | 139.41 | ||
| Claude 3.5 Sonnet (October 2024) | claude-3-5-sonnet-20241022 | Claude 3.5 Sonnet (Oct 2024) | Anthropic | United States of America | 2024-10-22 | API access | 2024-10-22 | Unverified | An updated version of the top model in the Claude 3.5 series, a previous flagship model series from Anthropic. | Claude 3.5 Sonnet (Oct 2024) | claude-3-5-sonnet-october-2024 | 134.12 | 128.60 | 138.94 | |
| Claude 3.5 Sonnet | claude-3-5-sonnet-20240620 | Claude 3.5 Sonnet (Jun 2024) | Anthropic | United States of America | 2024-06-20 | API access | 2024-06-20 | 2.7e+25 | Speculative | The first version of the top model in the Claude 3.5 series, a previous flagship model series from Anthropic. | Claude 3.5 Sonnet (Jun 2024) | claude-3-5-sonnet | 130.00 | ||
| Claude Sonnet 4 | claude-sonnet-4-20250514_59K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 59,000 reasoning tokens. | Claude Sonnet 4 (59k thinking) | claude-sonnet-4 | 142.30 | 139.81 | 144.40 | ||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_64K | Claude 3.7 Sonnet (64k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 64,000 reasoning tokens allowed. | Claude 3.7 Sonnet (64k thinking) | claude-3-7-sonnet | 141.64 | 138.97 | 143.73 |
| Grok 3 | grok-3-beta | xAI | United States of America | 2025-04-09 | API access | 2025-02-17 | 3.5e+26 | Likely | A beta release of XAI's third generation Grok model. | Grok 3 (beta) | grok-3 | 138.74 | 135.43 | 140.06 | |
| Magistral Small 1.1 | magistral-small-2506 | Mistral AI | France | 2025-06-10 | Open weights (unrestricted) | 2025-06-10 | Confident | magistral-small-1-1 | |||||||
| Qwen3-235B-A22B | qwen3-235b-a22b | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Likely | Qwen3 (Apr 2025) | qwen3-235b-a22b | 139.35 | 134.19 | 141.30 | ||
| Gemini 2.5 Pro (May 2025) | gemini-2.5-pro-preview-05-06 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-05-06 | API access | 2025-05-06 | Unknown | A May 2025 preview version of Gemini 2.5 Pro, a flagship reasoning model from Google DeepMind. | Gemini 2.5 Pro Preview (Jun 2025) | gemini-2-5-pro-may-2025 | 142.55 | 139.19 | 145.23 | |
| DeepSeek-R1 | DeepSeek-R1 | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-20 | 4.0e+24 | Confident | DeepSeek-R1 | deepseek-r1 | 139.33 | 136.87 | 141.20 | ||
| Claude Sonnet 4 | claude-sonnet-4-20250514_32K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 32,000 reasoning tokens. | Claude Sonnet 4 (32k thinking) | claude-sonnet-4 | 142.30 | 139.81 | 144.40 | ||
| Claude Opus 4 | claude-opus-4-20250514_16K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 16,000 reasoning tokens. | Claude Opus 4 (16k thinking) | claude-opus-4 | 143.06 | 140.01 | 145.22 | |
| Claude Sonnet 4 | claude-sonnet-4-20250514_16K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 16,000 reasoning tokens. | Claude Sonnet 4 (16k thinking) | claude-sonnet-4 | 142.30 | 139.81 | 144.40 | ||
| Gemini 2.5 Flash (May 2025) | gemini-2.5-flash-preview-05-20 | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-05-20 | API access | 2025-05-20 | Unknown | The May 2025 preview version of Gemini 2.5 Flash, a small reasoning model from Google DeepMind's Gemini 2.5 series. | gemini-2-5-flash-may-2025 | 141.92 | 139.59 | 143.67 | |||
| Qwen Plus | qwen-plus-2025-04-28 | Alibaba | China | 2025-04-28 | API access | 2024-02-06 | Unknown | Qwen Plus (Apr 2025) | qwen-plus | 130.82 | 124.98 | 132.43 | |||
| o3-mini | o3-mini-2025-01-31_medium | o3-mini (medium) | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | Unknown | A smaller version of OpenAI's o3 reasoning model, evaluated with medium reasoning effort. | o3-mini | 141.00 | 138.02 | 142.82 | ||
| DeepSeek-V3 | DeepSeek-V3-0324 | DeepSeek-V3 (Mar 2025) | DeepSeek | China | 2025-03-24 | Open weights (restricted use) | 2024-12-24 | 3.4e+24 | Confident | An updated version of the DeepSeek V3 model. | DeepSeek-V3 (Mar 2025) | deepseek-v3 | 136.51 | 132.41 | 138.30 |
| Gemini 2.5 Pro (Mar 2025) | gemini-2.5-pro-preview-03-25 | Gemini 2.5 Pro Preview (Mar 2025) | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-04-09 | API access | 2025-03-25 | Unknown | A March 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship model. | Gemini 2.5 Pro Preview (Mar 2025) | gemini-2-5-pro-mar-2025 | 144.03 | 141.78 | 146.09 | |
| Mistral Medium 3 | mistral-medium-2505 | Mistral AI | France | 2025-05-07 | API access | 2025-05-07 | Unknown | A 2025 language model from Mistral. | mistral-medium-3 | 134.88 | 124.93 | 136.49 | |||
| GPT-4.1 mini | gpt-4.1-mini-2025-04-14 | GPT-4.1 mini | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | Unknown | A smaller version of OpenAI's GPT-4.1. | gpt-4-1-mini | 135.28 | 130.61 | 137.08 | ||
| Gemini 2.0 Flash | gemini-2.0-flash-001 | Gemini 2.0 Flash (Feb 2025) | Google DeepMind,Google | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-02-05 | API access | 2024-12-11 | Unknown | The first version of Google DeepMind's Gemini 2.0 Flash model. | gemini-2-0-flash | 134.87 | 131.18 | 136.87 | ||
| Grok-3 mini | grok-3-mini-beta_low | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | Unknown | A beta release of a smaller version of XAI's third generation Grok model, evaluated with low reasoning effort. | grok-3-mini | 140.70 | 137.75 | 142.61 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219 | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic. | Claude 3.7 Sonnet | claude-3-7-sonnet | 141.64 | 138.97 | 143.73 | |
| Gemini 2.5 Flash (Apr 2025) | gemini-2.5-flash-preview-04-17 | Gemini 2.5 Flash Preview (Apr 2025) | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-04-17 | API access | 2025-04-17 | Unknown | An April 2025 preview of Gemini 2.5 Flash, a small model in Google DeepMind's Gemini 2.5 series. | gemini-2-5-flash-apr-2025 | 139.67 | 135.33 | 143.12 | ||
| GPT-4.1 nano | gpt-4.1-nano-2025-04-14 | GPT-4.1 nano | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | Unknown | The smallest version of OpenAI's GPT-4.1. | gpt-4-1-nano | 130.49 | 121.49 | 133.14 | ||
| QWQ-Plus | qwq-plus | Alibaba | China | 2025-04-08 | API access | 2025-04-08 | Unknown | QwQ-Plus | qwq-plus | ||||||
| Llama 4 Maverick | Llama-4-Maverick-17B-128E-Instruct-FP8 | Llama 4 Maverick (FP8) | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | Likely | The 17 billion active (400 billion total)-parameter model in Meta's Llama 4 series, quantized to FP8. | llama-4-maverick | 132.69 | 126.70 | 134.91 | |
| Llama 4 Scout | Llama-4-Scout-17B-16E-Instruct | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 4.1e+24 | Likely | The 17 billion active (109 billion total)-parameter model in Meta's Llama 4 series. | llama-4-scout | 129.82 | 123.74 | 131.82 | ||
| Qwen-Turbo | qwen-turbo-2024-11-01 | Alibaba | China | 2024-11-01 | API access | 2024-02-06 | Unknown | Qwen Turbo | qwen-turbo | ||||||
| Qwen Plus | qwen-plus-2025-01-25 | Alibaba | China | 2025-01-25 | API access | 2024-02-06 | Unknown | Qwen Plus (Jan 2025) | qwen-plus | 130.82 | 124.98 | 132.43 | |||
| Qwen2.5-Max | qwen-max-2025-01-25 | Alibaba | China | 2025-01-25 | API access | 2025-01-28 | Unknown | Qwen2.5-Max | qwen2-5-max | 132.98 | 127.68 | 136.39 | |||
| Hermes 2 Theta Llama-3 70B | Hermes-2-Theta-Llama-3-70B | Nous Research,Arcee AI | United States of America | 2024-06-20 | Open weights (restricted use) | 2024-06-20 | Confident | Nous Research’s fine-tuned Llama-3 70B optimized for instruction following and chat. | hermes-2-theta-llama-3-70b | ||||||
| Gemini 2.5 Pro (Mar 2025) | gemini-2.5-pro-exp-03-25 | Gemini 2.5 Pro Exp (Mar 2025) | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-03-25 | API access | 2025-03-25 | Unknown | A March 2025 preview version of Gemini 2.5 Pro. | Gemini 2.5 Pro Exp (Mar 2025) | gemini-2-5-pro-mar-2025 | 144.03 | 141.78 | 146.09 | |
| Mistral Small 3 | mistral-small-2501 | Mistral AI | France | 2025-01-25 | Open weights (unrestricted) | 2025-01-30 | 1.2e+24 | Confident | mistral-small-3 | ||||||
| Mistral Small 3.1 | mistral-small-2503 | Mistral AI | France | 2025-03-17 | Open weights (unrestricted) | 2025-03-17 | Confident | An updated version of Mistral small, a 24 billion-parameter model from Mistral. | mistral-small-3-1 | ||||||
| Gemma 3 27B | gemma-3-27b-it | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-03-12 | Open weights (restricted use) | 2025-03-12 | 2.3e+24 | Confident | An instruction-tuned version of the 27 billion-parameter model Google DeepMind's Gemma 3 series. | gemma-3-27b | 130.72 | 123.89 | 133.03 | ||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_32K | Claude 3.7 Sonnet (32k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 32,000 reasoning tokens allowed. | Claude 3.7 Sonnet (32k thinking) | claude-3-7-sonnet | 141.64 | 138.97 | 143.73 |
| Gemini 1.5 Flash 8B | gemini-1.5-flash-8b-001 | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-10-03 | API access | 2024-05-10 | Confident | gemini-1-5-flash-8b | |||||||
| DeepSeek-R1-Distill-Qwen-14B | DeepSeek-R1-Distill-Qwen-14B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | Speculative | A model based on Qwen 14B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | deepseek-r1-distill-qwen-14b | ||||||
| DeepSeek-R1-Distill-Llama-70B | DeepSeek-R1-Distill-Llama-70B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | Speculative | A model based on LLaMA 70B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | deepseek-r1-distill-llama-70b | ||||||
| DeepSeek-V3 | DeepSeek-V3 | DeepSeek | China | 2024-12-26 | Open weights (restricted use) | 2024-12-24 | 3.4e+24 | Confident | DeepSeek’s 2024 mixture-of-experts model. | DeepSeek-V3 | deepseek-v3 | 136.51 | 132.41 | 138.30 | |
| Tulu 3 (Tülu 3) 70B | Llama-3.1-Tulu-3-70B-DPO | Allen Institute for AI,University of Washington | United States of America | 2024-11-21 | Open weights (restricted use) | 2024-11-21 | Confident | Tülu 3 70B | tulu-3-tlu-3-70b | ||||||
| Gemma 2 27B | gemma-2-27b-it | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-06-24 | Open weights (restricted use) | 2024-06-24 | 2.1e+24 | Confident | An instruction-optimized version of Google DeepMind's Gemma 2 27B. | gemma-2-27b | 122.16 | 114.94 | 124.82 | ||
| Gemma 2 9B | gemma-2-9b-it | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-06-24 | Open weights (restricted use) | 2024-06-24 | 4.3e+23 | Confident | gemma-2-9b | 119.38 | 110.67 | 123.29 | |||
| Claude 2.1 | claude-2.1 | Anthropic | United States of America | 2023-11-21 | API access | 2023-11-21 | Unknown | An updated version in Anthropic's Claude 2 series. | claude-2-1 | 117.73 | 110.03 | 122.25 | |||
| Claude 2 | claude-2.0 | Anthropic | United States of America | 2023-07-11 | API access | 2023-07-11 | 3.9e+24 | Speculative | Anthropic's second generation Claude model. | claude-2 | 119.06 | 111.23 | 125.78 | ||
| Gemini 2.0 Flash Thinking (Jan 2025) | gemini-2.0-flash-thinking-exp-01-21 | Gemini 2.0 Flash Thinking Exp | Google DeepMind,Google | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-01-21 | API access | 2025-01-21 | Unknown | A January 2025 experimental version of Gemini 2.0 Flash Thinking, a small reasoning model from Google DeepMind. | gemini-2-0-flash-thinking-jan-2025 | 135.94 | 128.93 | 138.12 | ||
| o1-preview | o1-preview-2024-09-12 | o1-preview | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | Unknown | The September 2024 preview version of OpenAI’s first reasoning model, o1. | o1-preview | 135.25 | 130.11 | 138.51 | ||
| Gemini 2.0 Pro | gemini-2.0-pro-exp-02-05 | Gemini 2.0 Pro Exp (Feb 2025) | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-02-05 | Hosted access (no API) | 2024-12-11 | Unknown | A February 2025 experimental version of Google DeepMind's previous flagship model, Gemini 2.0 Pro. | gemini-2-0-pro | 135.35 | 129.32 | 137.94 | ||
| o1 | o1-2024-12-17_high | o1 (high) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | Unknown | OpenAI's first reasoning model, evaluated with at the high reasoning effort level. | o1 | 142.28 | 140.02 | 143.92 | ||
| GPT-4o (Aug 2024) | gpt-4o-2024-08-06 | GPT-4o (Aug 2024) | OpenAI | United States of America | 2024-08-06 | API access | 2024-08-06 | Speculative | An updated version of GPT-4o, OpenAI's model that powered ChatGPT from mid-2024 to -2025. | gpt-4o-aug-2024 | 128.95 | 123.21 | 132.12 | ||
| Gemini 1.5 Flash (Sep 2024) | gemini-1.5-flash-002 | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-09-24 | API access | 2024-09-24 | Unknown | The second version of Google DeepMind's Gemini 1.5 Flash. | gemini-1-5-flash-sep-2024 | 130.01 | 121.84 | 132.01 | |||
| o1-mini | o1-mini-2024-09-12_high | o1-mini (high) | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | Unknown | A smaller version of OpenAI’s first reasoning model, o1, evaluated with high reasoning effort. | o1-mini | 136.44 | 131.12 | 138.00 | ||
| o1-mini | o1-mini-2024-09-12_medium | o1-mini (medium) | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | Unknown | A smaller version of OpenAI’s first reasoning model, o1, evaluated with medium reasoning effort. | o1-mini | 136.44 | 131.12 | 138.00 | ||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_16K | Claude 3.7 Sonnet (16k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 16,000 reasoning tokens allowed. | Claude 3.7 Sonnet (16k thinking) | claude-3-7-sonnet | 141.64 | 138.97 | 143.73 |
| Grok-2 | grok-2-1212 | xAI | United States of America | 2024-12-12 | API access | 2024-08-13 | 3.0e+25 | Confident | XAI's second generation Grok model. | grok-2 | 130.37 | 124.89 | 132.13 | ||
| Mistral Large 2 | mistral-large-2411 | Mistral AI | France | 2024-11-18 | Open weights (non-commercial) | 2024-07-24 | 2.1e+25 | Likely | mistral-large-2 | 128.25 | 121.54 | 130.68 | |||
| GPT-4.5 | gpt-4.5-preview-2025-02-27 | GPT-4.5 Preview (Feb 2025) | OpenAI | United States of America | 2025-02-27 | API access | 2025-02-27 | 3.8e+26 | Likely | The largest model in OpenAI’s GPT series. | gpt-4-5 | 137.17 | 133.70 | 139.18 | |
| GPT-4 Turbo (Apr 2024) | gpt-4-turbo-2024-04-09 | OpenAI | United States of America | 2024-04-09 | API access | 2024-04-09 | Unknown | The April 2024 version of GPT-4 Turbo, OpenAI's then-flagship language model. | gpt-4-turbo-apr-2024 | 127.24 | 121.47 | 129.19 | |||
| o1 | o1-2024-12-17_medium | o1 (medium) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | Unknown | OpenAI's first reasoning model, evaluated with at the medium reasoning effort level. | o1 | 142.28 | 140.02 | 143.92 | ||
| Llama 3-70B | Meta-Llama-3-70B-Instruct | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.9e+24 | Confident | An instruction-tuned version of the 70 billion-parameter model in Meta’s LLaMA 3 series. | llama-3-70b | 119.30 | 87.91 | 125.18 | ||
| Gemini 1.5 Pro | gemini-1.5-pro-001 | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-05-24 | API access | 2024-02-15 | Speculative | The first of two versions of a previous flagship model from Google DeepMind, Gemini 1.5 Pro. | gemini-1-5-pro | 132.33 | 127.71 | 134.12 | |||
| Llama 3.3 70B | Llama-3.3-70B-Instruct | Meta AI | United States of America | 2024-12-06 | Open weights (restricted use) | 2024-12-06 | 6.9e+24 | Confident | An instruction-tuned version of the 70 billion-parameter model in Meta's Llama 3.3 series. | llama-3-3-70b | 126.92 | 121.05 | 129.38 | ||
| Llama 3.2 90B | Llama-3.2-90B-Vision-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | Confident | An instruction-tuned version of the 90 billion-parameter vision model in Meta's Llama 3.3 series. | llama-3-2-90b | 125.28 | 118.76 | 127.61 | |||
| Llama 2-70B | Llama-2-70b-chat-hf | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | Confident | A chat-optimized version of the 70-billion parameter model in Meta's Llama 2 series. | llama-2-70b | 113.52 | 103.75 | 117.40 | ||
| Gemini 1.5 Pro | gemini-1.5-pro-002 | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-09-24 | API access | 2024-02-15 | Speculative | The first of two versions of a previous flagship model from Google DeepMind, Gemini 1.5 Pro. | gemini-1-5-pro | 132.33 | 127.71 | 134.12 | |||
| Qwen2.5-32B | qwen2.5-32b-instruct | Alibaba | China | 2024-09-17 | Open weights (unrestricted) | 2024-09-17 | 3.5e+24 | Confident | Qwen2.5-32B (instruct) | qwen2-5-32b | |||||
| Mistral Large | mistral-large-2402 | Mistral AI | France | 2024-02-26 | API access | 2024-02-26 | 1.1e+25 | Likely | mistral-large | 120.85 | 109.45 | 124.10 | |||
| Claude 3 Sonnet | claude-3-sonnet-20240229 | Anthropic | United States of America | 2024-02-29 | API access | 2024-03-04 | Unknown | The mid-sized model in Anthropic's Claude 3 model series. | Claude 3 Sonnet | claude-3-sonnet | 119.63 | 110.76 | 123.86 | ||
| Gemini 1.0 Pro | gemini-1.0-pro-001 | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-02-15 | API access | 2023-12-06 | Speculative | An earlier flagship model from Google DeepMind. | gemini-1-0-pro | 116.55 | 108.08 | 120.77 | |||
| Llama 3.1-8B | Llama-3.1-8B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 1.2e+24 | Likely | An instruction-tuned version of the 8 billion-parameter model in Meta’s LLaMA 3.1 series. | llama-3-1-8b | 115.22 | 95.36 | 121.50 | ||
| Llama 3.1-405B | Llama-3.1-405B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 3.8e+25 | Confident | An instruction-tuned version of the 405 billion-parameter model in Meta’s LLaMA 3.1 series. | llama-3-1-405b | 127.77 | 119.26 | 130.28 | ||
| Qwen2.5-72B | qwen2.5-72b-instruct | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | Confident | An instruction-tuned version of the 72 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-72B | qwen2-5-72b | 129.26 | 122.12 | 131.40 | |
| GPT-4o (May 2024) | gpt-4o-2024-05-13 | GPT-4o (May 2024) | OpenAI | United States of America | 2024-05-13 | API access | 2024-05-13 | Speculative | The first version of GPT-4o, OpenAI's last multimodal model in the GPT-4 series, which powered ChatGPT from mid-2024 to -2025. | gpt-4o-may-2024 | 128.46 | 122.62 | 130.90 | ||
| Llama 3.1-70B | Llama-3.1-70B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 7.9e+24 | Confident | An instruction-optimized version of Llama 3.1 70B, the 70 billion-parameter model in Meta's Llama 3.1 series. | llama-3-1-70b | 125.05 | 117.47 | 126.98 | ||
| Gemini 1.5 Flash (May 2024) | gemini-1.5-flash-001 | Gemini 1.5 Flash (May 2024) | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-05-23 | API access | 2024-05-10 | Unknown | The first version of Google DeepMind's Gemini 1.5 Flash. | gemini-1-5-flash-may-2024 | 122.38 | 115.04 | 125.34 | ||
| Claude 3 Haiku | claude-3-haiku-20240307 | Anthropic | United States of America | 2024-03-07 | API access | 2024-03-04 | Unknown | The smallest model in Anthropic's Claude 3 series. | Claude 3 Haiku | claude-3-haiku | 117.19 | 106.98 | 121.33 | ||
| Llama 3-8B | Meta-Llama-3-8B-Instruct | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.2e+23 | Confident | An instruction-tuned version of the 8 billion-parameter model in Meta's Llama 3 series. | llama-3-8b | 116.00 | 105.44 | 119.81 | ||
| Claude 3 Opus | claude-3-opus-20240229 | Anthropic | United States of America | 2024-02-29 | API access | 2024-03-04 | Speculative | The largest model in Anthropic's Claude 3 model series. | Claude 3 Opus | claude-3-opus | 126.57 | 120.32 | 130.41 | ||
| Phi-4 | phi-4 | Microsoft Research | United States of America | 2024-12-12 | Open weights (unrestricted) | 2024-12-12 | 9.3e+23 | Confident | phi-4 | 130.76 | 119.41 | 132.90 | |||
| Mistral Large 2 | mistral-large-2407 | Mistral AI | France | 2024-07-24 | Open weights (non-commercial) | 2024-07-24 | 2.1e+25 | Likely | mistral-large-2 | 128.25 | 121.54 | 130.68 | |||
| phi-3-medium 14B | Phi-3-medium-128k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 4.0e+23 | Likely | phi-3-medium-14b | 121.16 | 113.93 | 126.03 | |||
| GPT-4 Turbo (Nov 2023) | gpt-4-0125-preview | GPT-4 Turbo Preview (January 2024) | OpenAI | United States of America | 2024-01-25 | API access | 2023-11-06 | Unknown | A January 2024 preview version of GPT-4 Turbo, OpenAI's updated GPT-4 series model. | gpt-4-turbo-nov-2023 | |||||
| GPT-3.5 Turbo | gpt-3.5-turbo-1106 | GPT-3.5 Turbo (Nov 2023) | OpenAI | United States of America | 2023-11-06 | API access | 2022-11-30 | Speculative | A November 2023 preview of GPT-3.5 turbo, OpenAI's updated GPT-3.5 series model. | gpt-3-5-turbo | 116.62 | 105.78 | 121.73 | ||
| GPT-4 Turbo (Nov 2023) | gpt-4-1106-preview | GPT-4 Turbo Preview (Nov 2023) | OpenAI | United States of America | 2023-11-06 | API access | 2023-11-06 | Unknown | A November 2023 preview of GPT-4 turbo, OpenAI's updated GPT-4 series model. | gpt-4-turbo-nov-2023 | |||||
| GPT-3.5 Turbo | gpt-3.5-turbo-0125 | GPT-3.5 Turbo (Jan 2024) | OpenAI | United States of America | 2024-01-25 | API access | 2022-11-30 | Speculative | A January 2024 preview version of GPT-3.5 Turbo, OpenAI's updated GPT-3.5 series model. | gpt-3-5-turbo | 116.62 | 105.78 | 121.73 | ||
| Yi-1.5-34B | Yi-1.5-34B-Chat | 01.AI | China | 2024-05-13 | Open weights (restricted use) | 2024-05-13 | 7.3e+23 | Confident | Yi-1.5-34B (chat) | yi-1-5-34b | |||||
| Yi-34B | Yi-34B-Chat | 01.AI | China | 2023-11-22 | Open weights (restricted use) | 2023-11-02 | 6.1e+23 | Confident | Yi-34B (chat) | yi-34b | 117.26 | 108.22 | 122.34 | ||
| Qwen2-72B | qwen2-72b-instruct | Alibaba | China | 2024-06-07 | Open weights (unrestricted) | 2024-06-07 | 3.0e+24 | Confident | Qwen's 72 billion-parameter model in the Qwen 2 series. | Qwen2-72B | qwen2-72b | 125.31 | 116.82 | 127.35 | |
| Qwen1.5-72B | qwen1.5-72b-chat | Alibaba | China | 2024-02-04 | Open weights (restricted use) | 2024-02-04 | 1.3e+24 | Confident | Qwen1.5-72B | qwen1-5-72b | |||||
| Qwen1.5-32B | qwen1.5-32b-chat | Alibaba | China | 2024-04-03 | Open weights (restricted use) | 2024-02-05 | Confident | A chat-optimized 32 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-32B (chat) | qwen1-5-32b | |||||
| Mistral 7B | Mistral-7B-Instruct-v0.3 | Mistral AI | France | 2024-05-27 | Open weights (unrestricted) | 2023-10-10 | Confident | An instruction-tuned version of Mistral-7B-v0.3, an updated version of their 7 billion-parameter model. | mistral-7b | 110.69 | 100.53 | 116.26 | |||
| DeepSeek LLM 67B | deepseek-llm-67b-chat | DeepSeek | China | 2023-11-29 | Open weights (restricted use) | 2024-01-05 | 8.0e+23 | Confident | DeepSeek LLM 67B (chat) | deepseek-llm-67b | |||||
| Mixtral 8x7B | Mixtral-8x7B-Instruct-v0.1 | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | Speculative | An instruction-tuned version of Mistral's 7 billion active (56 billion total)-parameter mixture-of-experts model | mixtral-8x7b | 118.00 | 108.69 | 122.19 | ||
| WizardLM-2 8x22B | WizardLM-2-8x22B | Microsoft | United States of America | 2024-04-15 | Open weights (unrestricted) | 2024-04-15 | Confident | WizardLM-2 8x22B | wizardlm-2-8x22b | ||||||
| DBRX | dbrx-instruct | Databricks | United States of America | 2024-03-27 | Open weights (restricted use) | 2024-03-27 | 2.6e+24 | Confident | DBRX (instruct) | dbrx | |||||
| Ministral 3B | ministral-3b-2410 | Mistral AI | France | 2024-10-16 | API access | 2024-10-16 | Confident | ministral-3b | |||||||
| Mixtral 8x22B | open-mixtral-8x22b | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | Speculative | mixtral-8x22b | 120.88 | 106.97 | 124.40 | |||
| Ministral 8B | ministral-8b-2410 | Mistral AI | France | 2024-10-16 | Open weights (non-commercial) | 2024-10-16 | Confident | ministral-8b | |||||||
| Mixtral 8x7B | open-mixtral-8x7b | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | Speculative | mixtral-8x7b | 118.00 | 108.69 | 122.19 | |||
| Mistral 7B | open-mistral-7b | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | Confident | mistral-7b | 110.69 | 100.53 | 116.26 | ||||
| Mistral NeMo | open-mistral-nemo-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | Confident | mistral-nemo | 117.94 | 109.61 | 121.82 | ||||
| Eurus-2-7B-PRIME | Eurus-2-7B-PRIME | Tsinghua University,University of Illinois Urbana-Champaign (UIUC),Shanghai AI Lab,Peking University,Shanghai Jiao Tong University,CUHK Shenzhen Research Institute | China,United States of America | 2024-12-31 | Open weights (unrestricted) | 2025-02-03 | Speculative | eurus-2-7b-prime | |||||||
| Gemini 2.5 Deep Think | gemini-2.5-deep-think-2025-08-01-webapp | Google,Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-08-01 | Hosted access (no API) | 2025-08-01 | Unknown | gemini-2-5-deep-think | |||||||
| MPT-7B | mpt-7b | MosaicML | United States of America | 2023-05-05 | Open weights (unrestricted) | 2023-05-05 | 4.2e+22 | Confident | mpt-7b | 88.80 | 68.94 | 97.58 | |||
| chatglm2-6b | 2023-06-24 | Unverified | chatglm2-6b | ||||||||||||
| internlm-7b | 2023-07-05 | Unverified | internlm-7b | ||||||||||||
| internlm-20b | 2023-09-18 | Unverified | internlm-20b | ||||||||||||
| Baichuan 2-7B | Baichuan-2-7B-Base | Baichuan | China | 2023-09-20 | Open weights (restricted use) | 2023-09-20 | 1.1e+23 | Confident | baichuan-2-7b | 92.28 | 76.74 | 103.61 | |||
| Baichuan2-13B | Baichuan-2-13B-Base | Baichuan | China | 2023-09-06 | Open weights (restricted use) | 2023-09-06 | 2.0e+23 | Confident | baichuan2-13b | 100.18 | 79.47 | 108.45 | |||
| LLaMA-7B | LLaMA-7B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-24 | 4.0e+22 | Confident | Meta's smallest model in the original Llama series. | llama-7b | 91.57 | 76.35 | 100.47 | ||
| LLaMA-13B | LLaMA-13B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-27 | 7.8e+22 | Confident | Meta's 13 billion-parameter model in the original Llama series. | llama-13b | 96.70 | 81.54 | 104.05 | ||
| LLaMA-33B | LLaMA-33B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-27 | 2.7e+23 | Confident | Meta's 33 billion-parameter model in the original Llama series | llama-33b | 105.42 | 92.92 | 111.62 | ||
| LLaMA-65B | LLaMA-65B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-24 | 5.5e+23 | Confident | Meta's largest model in the original Llama series. | llama-65b | 108.80 | 98.29 | 114.31 | ||
| Llama 2-7B | Llama-2-7b | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.4e+22 | Confident | Meta's 7 billion-parameter model in the Llama 2 series of open source models. | llama-2-7b | 94.63 | 80.01 | 102.17 | ||
| Llama 2-13B | Llama-2-13b | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 1.6e+23 | Confident | Meta's 13 billion-parameter model in the Llama 2 series of open source models. | llama-2-13b | 103.72 | 92.02 | 110.35 | ||
| Llama 2-70B | Llama-2-70b-hf | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | Confident | A chat-optimized of a 70 billion-parameter model in the Llama 2 series. | llama-2-70b | 113.52 | 103.75 | 117.40 | ||
| Stable Beluga 2 | StableBeluga2 | Stability AI | United Kingdom of Great Britain and Northern Ireland | 2023-07-20 | Open weights (non-commercial) | 2023-07-20 | Likely | Stable Beluga 2 | stable-beluga-2 | 116.41 | 107.21 | 120.60 | |||
| Qwen-1_8B | Qwen-1.8B | 2023-11-30 | Unverified | qwen-1-8b | |||||||||||
| Qwen-7B | Qwen-7B | Alibaba | China | 2023-09-28 | Open weights (restricted use) | 2023-09-28 | 1.0e+23 | Confident | Qwen's 7 billion-parameter model in the original Qwen series. | Qwen-7B | qwen-7b | 104.81 | 85.41 | 111.58 | |
| Qwen-14B | Qwen-14B | Alibaba | China | 2023-09-28 | Open weights (restricted use) | 2023-09-28 | 2.5e+23 | Confident | Qwen's 14 billion-parameter model in the original Qwen series. | Qwen-14B | qwen-14b | 111.82 | 99.59 | 117.97 | |
| InstructGPT 175B | text-davinci-001 | OpenAI | United States of America | 2022-01-27 | API access | 2022-01-27 | 3.2e+23 | Confident | instructgpt-175b | ||||||
| Gopher (280B) | Gopher (280B) | DeepMind | United Kingdom of Great Britain and Northern Ireland | 2021-12-08 | Unreleased | 2021-12-08 | 6.3e+23 | Confident | A 280 billion-parameter model from DeepMind in 2021. | gopher-280b | |||||
| Megatron-Turing NLG 530B | Megatron-Turing NLG 530B | Microsoft,NVIDIA | United States of America | 2022-01-28 | Unreleased | 2021-10-11 | 8.6e+23 | Confident | megatron-turing-nlg-530b | ||||||
| Chinchilla | Chinchilla (70B) | DeepMind | United Kingdom of Great Britain and Northern Ireland | 2022-03-29 | Unreleased | 2022-03-29 | 5.8e+23 | Confident | Chinchilla | chinchilla | |||||
| PaLM (540B) | PaLM 540B | Google Research | United States of America | 2022-04-04 | Unreleased | 2022-04-04 | 2.5e+24 | Confident | Google's largest model in the PaLM series. | palm-540b | |||||
| Inflection-1 | Inflection-1 | Inflection AI | United States of America | 2023-06-22 | Hosted access (no API) | 2023-06-23 | 1.0e+24 | Speculative | inflection-1 | ||||||
| Falcon-7B | falcon-7b | Technology Innovation Institute | United Arab Emirates | 2023-04-24 | Open weights (unrestricted) | 2023-04-24 | 6.3e+22 | Confident | The 7 billion-parameter model in the Falcon series. | falcon-7b | 89.49 | 75.88 | 101.73 | ||
| Falcon-40B | falcon-40b | Technology Innovation Institute | United Arab Emirates | 2023-03-15 | Open weights (unrestricted) | 2023-03-15 | 2.4e+23 | Confident | The 40 billion-parameter model in the Falcon series. | falcon-40b | 101.71 | 89.28 | 109.80 | ||
| Falcon-180B | falcon-180B | Technology Innovation Institute | United Arab Emirates | 2023-09-06 | Open weights (restricted use) | 2023-09-06 | 3.8e+24 | Confident | The 180 billion-parameter model in the Falcon series. | falcon-180b | 111.12 | 102.28 | 128.34 | ||
| Amazon Nova Pro | amazon.nova-pro-v1:0 | Amazon | United States of America | 2024-12-03 | API access | 2024-12-03 | 6.0e+24 | Speculative | amazon-nova-pro | ||||||
| Gemini 2.5 Flash (Sep 2025) | gemini-2.5-flash-preview-09-2025 | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-09-25 | API access | 2025-09-25 | Unknown | Gemini 2.5 Flash (Sep 2025) | gemini-2-5-flash-sep-2025 | 142.11 | 137.83 | 144.32 | |||
| DeepSeek-V3.2-Exp | DeepSeek-V3.2-Exp | DeepSeek | China | 2025-09-29 | Open weights (unrestricted) | 2025-09-29 | 3.8e+24 | Confident | deepseek-v3-2-exp | 144.66 | 141.39 | 148.55 | |||
| Mistral 7B | Mistral-7B-Instruct-v0.2 | Mistral AI | France | Open weights (unrestricted) | 2023-10-10 | Confident | mistral-7b | 110.69 | 100.53 | 116.26 | |||||
| Gemma 2B | gemma-2b | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-02-21 | Open weights (restricted use) | 2024-02-21 | 4.5e+22 | Confident | Google's 2 billion-parameter model in the Gemma series of open source models. | gemma-2b | 89.69 | 73.23 | 98.74 | ||
| Gemma 7B | gemma-7b | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-02-21 | Open weights (restricted use) | 2024-02-21 | 3.1e+23 | Confident | Google's 7 billion-parameter model in the Gemma series of open source models. | gemma-7b | 110.42 | 101.35 | 116.94 | ||
| Baichuan1-7B | Baichuan-7B | Baichuan | China | 2023-06-01 | Open weights (non-commercial) | 2023-06-01 | 5.0e+22 | Confident | baichuan1-7b | 85.57 | 67.37 | 96.26 | |||
| phi-3.5-mini | Phi-3.5-mini-instruct | Microsoft | United States of America | 2024-08-16 | Open weights (unrestricted) | 2024-04-23 | 3.7e+22 | Confident | An instruction-tuned version of phi-3.5-mini, a small model in Microsoft's Phi 3.5 open source model series. | phi-3-5-mini | |||||
| Phi-3.5-MoE | Phi-3.5-MoE-instruct | Microsoft | United States of America | 2024-08-17 | Open weights (unrestricted) | 2024-04-23 | 3.0e+23 | Confident | phi-3-5-moe | ||||||
| Mistral 7B | Mistral-7B-v0.1 | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | Confident | mistral-7b | 110.69 | 100.53 | 116.26 | ||||
| Mistral NeMo | Mistral-Nemo-Base-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | Confident | mistral-nemo | 117.94 | 109.61 | 121.82 | ||||
| Gemma 2 9B | gemma-2-9b | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | Open weights (restricted use) | 2024-06-24 | 4.3e+23 | Confident | A 9 billion-parameter open source model from Google DeepMind in the Gemma 2 series. | gemma-2-9b | 119.38 | 110.67 | 123.29 | |||
| DeepSeek-V2 (MoE-236B) | DeepSeek-V2 | DeepSeek | China | 2024-05-07 | Open weights (restricted use) | 2024-05-07 | 1.0e+24 | Confident | deepseek-v2-moe-236b | 125.54 | 118.21 | 128.27 | |||
| Qwen2.5-72B | Qwen2.5-72B | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | Confident | A 72 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-72B | qwen2-5-72b | 129.26 | 122.12 | 131.40 | |
| Llama 3.1-405B | Llama-3.1-405B | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 3.8e+25 | Confident | The 405 billion-parameter model in Meta’s LLaMA 3.1 series. | llama-3-1-405b | 127.77 | 119.26 | 130.28 | ||
| DeepSeek-V3 | DeepSeek-V3-Base | DeepSeek | China | 2024-12-26 | Open weights (restricted use) | 2024-12-24 | 3.4e+24 | Confident | DeepSeek-V3 (base) | deepseek-v3 | 136.51 | 132.41 | 138.30 | ||
| MPT-30B | mpt-30b | MosaicML | United States of America | 2023-06-22 | Open weights (unrestricted) | 2023-06-22 | 1.9e+23 | Confident | mpt-30b | 97.24 | 82.73 | 104.72 | |||
| Llama 2-34B | Llama-2-34b | Meta AI | United States of America | 2023-07-18 | Unreleased | 2023-07-18 | 4.1e+23 | Confident | Meta's 34 billion-parameter model in the Llama 2 series. | llama-2-34b | 102.71 | 89.35 | 109.09 | ||
| PaLM 62B | 2022-04-04 | Unverified | palm-62b | ||||||||||||
| Mistral 7B | Mistral-7B-Instruct-v0.1 | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | Confident | mistral-7b | 110.69 | 100.53 | 116.26 | ||||
| Mixtral 8x7B | Mixtral-8x7B-v0.1 | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | Speculative | mixtral-8x7b | 118.00 | 108.69 | 122.19 | |||
| Nemotron-4 15B | Nemotron-4 15B | NVIDIA | United States of America | 2024-02-26 | Unreleased | 2024-02-27 | 7.5e+23 | Confident | A 15 billion-parameter language model trained by Nvidia. | nemotron-4-15b | 105.27 | 93.52 | 111.38 | ||
| vicuna-13b-v1.1 | 2023-04-12 | Unverified | vicuna-13b-v1-1 | ||||||||||||
| OPT-1.3B | opt-1.3b | Meta AI | United States of America | 2022-05-11 | Open weights (non-commercial) | 2022-06-21 | Confident | opt-1-3b | |||||||
| GPT-Neo-2.7B | gpt-neo-2.7B | EleutherAI | United States of America | 2023-03-30 | Open weights (unrestricted) | 2021-03-21 | 7.9e+21 | Confident | gpt-neo-2-7b | ||||||
| GPT-2 (1.5B) | gpt2-xl | OpenAI | United States of America | 2019-11-05 | Open weights (unrestricted) | 2019-02-14 | 1.9e+21 | Speculative | The largest version of GPT-2, OpenAI's second generation transformer-based large pretrained language model. | gpt-2-1-5b | |||||
| XGen-7B | xgen-7b-8k-base | Salesforce | United States of America | 2023-06-27 | Open weights (unrestricted) | 2023-09-07 | 8.0e+22 | Confident | XGen-7B | xgen-7b | 86.72 | 66.70 | 96.26 | ||
| open_llama_7b | 2023-06-07 | Unverified | open-llama-7b | ||||||||||||
| RedPajama-INCITE-7B-Base | 2023-05-04 | Unverified | redpajama-incite-7b-base | ||||||||||||
| GPT-NeoX-20B | gpt-neox-20b | EleutherAI | United States of America | 2022-04-07 | Open weights (unrestricted) | 2022-02-09 | 9.3e+22 | Confident | gpt-neox-20b | ||||||
| opt-13b | 2022-05-11 | Unverified | opt-13b | ||||||||||||
| GPT-J-6B | gpt-j-6b | EleutherAI,LAION | United States of America,Germany | 2021-08-05 | Open weights (unrestricted) | 2021-05-01 | 1.5e+22 | Confident | gpt-j-6b | ||||||
| Dolly 2.0-12b | dolly-v2-12b | Databricks | United States of America | 2023-04-11 | Open weights (unrestricted) | 2023-04-12 | Confident | dolly-2-0-12b | 80.88 | 59.97 | 93.81 | ||||
| Cerebras-GPT-13B | Cerebras-GPT-13B | Cerebras Systems | United States of America | 2023-03-20 | Open weights (unrestricted) | 2023-04-06 | 2.3e+22 | Confident | cerebras-gpt-13b | 71.40 | 46.75 | 87.09 | |||
| stablelm-tuned-alpha-7b | 2023-04-19 | Unverified | stablelm-tuned-alpha-7b | ||||||||||||
| Yi 6B | Yi-6B | 01.AI | China | 2023-11-02 | Open weights (restricted use) | 2023-11-02 | 1.3e+23 | Confident | Yi-6B (chat) | yi-6b | 102.09 | 89.88 | 109.56 | ||
| Yi-34B | Yi-34B | 01.AI | China | 2023-11-02 | Open weights (restricted use) | 2023-11-02 | 6.1e+23 | Confident | Yi-34B | yi-34b | 117.26 | 108.22 | 122.34 | ||
| GPT-3.5 Turbo | gpt-3.5-turbo-0613 | GPT-3.5 Turbo (June 2023) | OpenAI | United States of America | 2023-06-13 | API access | 2022-11-30 | Speculative | GPT-3.5 Turbo (June 2023) | gpt-3-5-turbo | 116.62 | 105.78 | 121.73 | ||
| Baichuan 1-13B | Baichuan-13B-Base | Baichuan | China | 2023-07-11 | Open weights (restricted use) | 2023-07-11 | 9.4e+22 | Confident | baichuan-1-13b | ||||||
| phi-3-mini 3.8B | Phi-3-mini-4k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 7.5e+22 | Confident | A 4 billion-parameter model in Microsoft's Phi-3 open source model series with a 4000-token context length. | phi-3-mini-3-8b | 117.03 | 103.81 | 123.67 | ||
| phi-3-small 7.4B | Phi-3-small-8k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 2.1e+23 | Confident | A 7 billion-parameter model in Microsoft's Phi-3 open source model series with a 8000-token context length. | phi-3-small-7-4b | 122.88 | 115.09 | 129.62 | ||
| Phi-2 | phi-2 | Microsoft | United States of America | 2023-12-12 | Open weights (unrestricted) | 2023-12-12 | 2.3e+22 | Confident | phi-2 | 105.54 | 55.85 | 112.79 | |||
| Llama 2-13B | Llama-2-13b-chat | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 1.6e+23 | Confident | A chat-optimized version of Meta's 13 billion-parameter model in the Llama 2 series. | llama-2-13b | 103.72 | 92.02 | 110.35 | ||
| Llama 2-70B | Llama-2-70b-chat | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | Confident | A chat-optimized version of Meta's 70 billion-parameter model in the Llama 2 series. | llama-2-70b | 113.52 | 103.75 | 117.40 | ||
| Baichuan2-13B-Chat | 2023-09-06 | Unverified | baichuan2-13b-chat | ||||||||||||
| Qwen-14B | Qwen-14B-Chat | Alibaba | China | 2023-09-24 | Open weights (restricted use) | 2023-09-28 | 2.5e+23 | Confident | A chat-optimized version of Qwen's first-generation 14 billion-parameter model. | Qwen-14B (chat) | qwen-14b | 111.82 | 99.59 | 117.97 | |
| internlm-chat-20b | 2023-09-17 | Unverified | internlm-chat-20b | ||||||||||||
| Yi 6B | Yi-6B-Chat | 01.AI | China | 2023-11-22 | Open weights (restricted use) | 2023-11-02 | 1.3e+23 | Confident | Yi-6B | yi-6b | 102.09 | 89.88 | 109.56 | ||
| INTELLECT-1 | INTELLECT-1-Instruct | Prime Intellect,Hugging Face,Arcee AI | United States of America | 2024-11-29 | 2024-11-29 | 6.0e+22 | Confident | intellect-1 | |||||||
| Yi-9B | 2024-03-01 | Unverified | yi-9b | ||||||||||||
| PaLM 2-S | 2023-05-17 | Unverified | palm-2-s | ||||||||||||
| PaLM 2-M | 2023-05-17 | Unverified | palm-2-m | ||||||||||||
| PaLM 2-L | 2023-05-17 | Unverified | palm-2-l | ||||||||||||
| Qwen2.5-Coder-0.5B | 2024-09-18 | Unverified | qwen2-5-coder-0-5b | ||||||||||||
| DeepSeek Coder 1.3B | deepseek-coder-1.3b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 1.6e+22 | Likely | DeepSeek Coder 1.3B (base) | deepseek-coder-1-3b | 54.45 | 28.85 | 78.35 | ||
| Qwen2.5-Coder (1.5B) | Qwen2.5-Coder-1.5B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-09-18 | 5.1e+22 | Confident | Qwen's 1.5 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-1.5B | qwen2-5-coder-1-5b | 100.49 | 75.57 | 110.70 | |
| StarCoder 2 3B | starcoder2-3b | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | 2024-02-22 | Open weights (restricted use) | 2024-02-29 | 5.9e+22 | Confident | StarCoder 2 3B | starcoder-2-3b | 83.98 | 61.81 | 95.59 | ||
| Qwen2.5-Coder-3B | 2024-09-18 | Unverified | qwen2-5-coder-3b | ||||||||||||
| StarCoder 2 7B | starcoder2-7b | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | 2024-02-20 | Open weights (restricted use) | 2024-02-29 | 1.5e+23 | Confident | StarCoder 2 7B | starcoder-2-7b | 89.60 | 67.07 | 101.06 | ||
| DeepSeek Coder 6.7B | deepseek-coder-6.7b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 8.0e+22 | Likely | DeepSeek Coder 6.7B (base) | deepseek-coder-6-7b | 85.03 | 62.07 | 95.86 | ||
| DeepSeek-Coder-V2-Lite-Base | 2024-06-13 | Unverified | deepseek-coder-v2-lite-base | ||||||||||||
| CodeQwen1.5-7B | 2024-04-15 | Unverified | codeqwen1-5-7b | ||||||||||||
| Qwen2.5-Coder (7B) | Qwen2.5-Coder-7B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-09-18 | 2.5e+23 | Confident | Qwen's 7 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-7B | qwen2-5-coder-7b | 111.89 | 95.03 | 119.46 | |
| StarCoder 2 15B | starcoder2-15b | Hugging Face,ServiceNow,NVIDIA,BigCode | United States of America | 2024-02-20 | Open weights (restricted use) | 2024-02-29 | 3.9e+23 | Confident | StarCoder 2 15B | starcoder-2-15b | 102.57 | 86.11 | 112.31 | ||
| Qwen2.5-Coder-14B | 2024-09-18 | Unverified | qwen2-5-coder-14b | ||||||||||||
| DeepSeek Coder 33B | deepseek-coder-33b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 4.0e+23 | Likely | DeepSeek Coder 33B (base) | deepseek-coder-33b | 92.82 | 72.52 | 102.82 | ||
| DeepSeek-Coder-V2 236B | DeepSeek-Coder-V2-Base | DeepSeek | China | 2024-06-17 | Open weights (restricted use) | 2024-06-17 | 1.3e+24 | Confident | DeepSeek-Coder-V2 (base) | deepseek-coder-v2-236b | |||||
| Qwen2.5-Coder (32B) | Qwen2.5-Coder-32B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-11-12 | 1.1e+24 | Confident | Qwen's 32 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-32B | qwen2-5-coder-32b | 121.04 | 108.50 | 128.22 | |
| Phi-1.5 | phi-1_5 | Microsoft | United States of America | 2023-09-11 | Open weights (unrestricted) | 2023-09-11 | 1.2e+21 | Confident | phi-1-5 | 85.13 | 35.73 | 99.48 | |||
| GPT-3.5 | text-davinci-002 | OpenAI | United States of America | 2022-03-15 | API access | 2022-11-28 | 2.6e+24 | Speculative | gpt-3-5 | ||||||
| Claude Instant | claude-instant-1.1 | Anthropic | United States of America | API access | 2023-08-09 | Unknown | claude-instant | 121.11 | 111.41 | 125.00 | |||||
| Claude Instant | claude-instant-1.2 | Anthropic | United States of America | 2023-08-09 | API access | 2023-08-09 | Unknown | claude-instant | 121.11 | 111.41 | 125.00 | ||||
| T5-Large | Unverified | t5-large | |||||||||||||
| T5-11B | T5-11B | United States of America | Open weights (unrestricted) | 2019-10-23 | 3.3e+22 | Confident | The largest, 11 billion-parameter version of T5, an early large language model from Google. | t5-11b | |||||||
| T5-3B | T5-3B | United States of America | Open weights (unrestricted) | 2019-10-23 | 9.0e+21 | Confident | The 3 billion-parameter version of T5, an early large language model from Google. | t5-3b | |||||||
| Mixtral 8x22B | Mixtral-8x22B-Instruct-v0.1 | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | Speculative | An instruction-tuned version of Mistral's 8 billion active (176 billion total)-parameter mixture-of-experts model | mixtral-8x22b | 120.88 | 106.97 | 124.40 | ||
| Qwen2.5-72B | Qwen2.5-VL-72B-Instruct | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | Confident | A 72 billion-parameter instruction-tuned vision language model in the Qwen 2.5 series. | Qwen2.5-VL-72B | qwen2-5-72b | 129.26 | 122.12 | 131.40 | |
| Qwen2-VL-72B-Instruct | 2024-08-29 | Unverified | qwen2-vl-72b-instruct | ||||||||||||
| LLaVA-Video-72B-Qwen2 | 2024-09-02 | Unverified | llava-video-72b-qwen2 | ||||||||||||
| Oryx-1.5-32B | 2024-10-22 | Unverified | oryx-1-5-32b | ||||||||||||
| InternVL2_5-78B | InternVL2_5-78B | Shanghai AI Lab,SenseTime,Tsinghua University,Nanjing University,Fudan University,Chinese University of Hong Kong (CUHK),Shanghai Jiao Tong University | China,Hong Kong | 2024-12-06 | Open weights (restricted use) | 2024-12-06 | Confident | internvl2-5-78b | |||||||
| ViLAMP-llava-qwen | 2025-05-01 | Unverified | vilamp-llava-qwen | ||||||||||||
| Aria | 2024-09-30 | Unverified | aria | ||||||||||||
| LinVT | 2024-12-06 | Unverified | linvt | ||||||||||||
| LLaVA-Video-7B-Qwen2-TPO | 2025-01-19 | Unverified | llava-video-7b-qwen2-tpo | ||||||||||||
| VideoLLAMA3-7B | 2024-01-22 | Unverified | videollama3-7b | ||||||||||||
| LiveCC-7B-Instruct | 2024-04-12 | Unverified | livecc-7b-instruct | ||||||||||||
| LLaVA-Video-7B-Qwen2 | 2024-09-02 | Unverified | llava-video-7b-qwen2 | ||||||||||||
| NVILA 8B | NVILA-8B | NVIDIA,Massachusetts Institute of Technology (MIT),University of California (UC) Berkeley,University of California San Diego,University of Washington,Tsinghua University | China,United States of America | 2024-12-10 | Open weights (non-commercial) | 2024-12-05 | 2.3e+21 | Likely | nvila-8b | ||||||
| VideoChat-Flash-Qwen2-7B_res448 | 2025-01-11 | Unverified | videochat-flash-qwen2-7b-res448 | ||||||||||||
| LLaVA-OneVision 72B | 2024-08-06 | Unverified | llava-onevision-72b | ||||||||||||
| Qwen2-VL-7B-Instruct | 2024-08-29 | Unverified | qwen2-vl-7b-instruct | ||||||||||||
| ByteVideoLLM-14B | 2024-10-13 | Unverified | bytevideollm-14b | ||||||||||||
| mPLUG-Owl3-7B-241101 | 2024-11-26 | Unverified | mplug-owl3-7b-241101 | ||||||||||||
| MiniCPM-o-2_6 | 2025-01-12 | Unverified | minicpm-o-2-6 | ||||||||||||
| VideoLLAMA2-7B | 2024-06-12 | Unverified | videollama2-7b | ||||||||||||
| MiniCPM-V-2_6 | 2024-08-03 | Unverified | minicpm-v-2-6 | ||||||||||||
| GPT-4V | gpt-4-1106-vision-preview | OpenAI | United States of America | 2023-11-06 | API access | 2023-09-25 | Unknown | gpt-4v | |||||||
| TimeMarker | 2024-11-27 | Unverified | timemarker | ||||||||||||
| InternVL2-40B | InternVL2-40B | Shanghai AI Lab | China | 2024-07-08 | Open weights (unrestricted) | 2024-07-04 | Confident | internvl2-40b | |||||||
| Video-XL-7B | 2024-10-17 | Unverified | video-xl-7b | ||||||||||||
| VITA | 2024-08-12 | Unverified | vita | ||||||||||||
| VITA-1.5 | 2024-12-20 | Unverified | vita-1-5 | ||||||||||||
| kangaroo | 2024-11-13 | Unverified | kangaroo | ||||||||||||
| Video-CCAM-7B-v1.2 | 2024-09-29 | Unverified | video-ccam-7b-v1-2 | ||||||||||||
| long-llava-qwen2-7b | 2024-08-30 | Unverified | long-llava-qwen2-7b | ||||||||||||
| LongVA-7B | 2024-06-13 | Unverified | longva-7b | ||||||||||||
| InternVL-Chat-V1-5 | 2024-04-18 | Unverified | internvl-chat-v1-5 | ||||||||||||
| Qwen-VL-Max | 2024-01-18 | Unverified | qwen-vl-max | ||||||||||||
| Chat-Uni-Vi-7B-v1.5 + 100k SG-WV | 2024-06-20 | Unverified | chat-uni-vi-7b-v1-5--100k-sg-wv | ||||||||||||
| SliME-Llama3-8B | 2024-06-02 | Unverified | slime-llama3-8b | ||||||||||||
| Chat-UniVi-7B-v1.5 | 2024-04-23 | Unverified | chat-univi-7b-v1-5- | ||||||||||||
| video_chat2_mistral | 2023-11-29 | Unverified | video-chat2-mistral | ||||||||||||
| sharegpt4video-8b | 2024-05-27 | Unverified | sharegpt4video-8b | ||||||||||||
| ST-LLM | 2024-03-28 | Unverified | st-llm | ||||||||||||
| Qwen-VL-Chat | 2023-08-20 | Unverified | qwen-vl-chat | ||||||||||||
| Video-LLaVA-7B | 2023-11-17 | Unverified | video-llava-7b | ||||||||||||
| Qwen3-Coder-480B-A35B | Qwen3-Coder-480B-A35B-Instruct | Alibaba | China | 2025-07-31 | Open weights (unrestricted) | 2025-07-22 | 1.6e+24 | Confident | Qwen3 Coder | qwen3-coder-480b-a35b | 142.88 | 138.88 | 144.93 | ||
| Kimi K2 | Kimi-K2-Instruct | Kimi K2 Instruct | Moonshot | China | 2025-07-12 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | Moonshot AI’s 1-trillion parameter scale model. | kimi-k2 | 140.37 | 137.40 | 142.49 | |
| gpt-oss-120b | gpt-oss-120b | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | The larger one of OpenAI's two latest open source language models. | gpt-oss-120b | 139.64 | 135.24 | 143.10 | ||
| Llama 4 Maverick | Llama-4-Maverick-17B-128E-Instruct | Llama 4 Maverick | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | Likely | The 17 billion active (400 billion total)-parameter model in Meta's Llama 4 series. | llama-4-maverick | 132.69 | 126.70 | 134.91 | |
| Qwen2.5-Coder (32B) | Qwen2.5-Coder-32B-Instruct | Alibaba | China | 2024-11-21 | Open weights (unrestricted) | 2024-11-12 | 1.1e+24 | Confident | An instruction-tuned and coding-optimized 32 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-Coder-32B (instruct) | qwen2-5-coder-32b | 121.04 | 108.50 | 128.22 | |
| GLaM | GLaM (MoE) | United States of America | 2021-12-13 | Unreleased | 2021-12-13 | 3.6e+23 | Confident | glam | |||||||
| Amazon Nova Lite | amazon.nova-lite-v1:0 | Amazon | United States of America | 2024-12-03 | API access | 2024-12-03 | Unknown | amazon-nova-lite | |||||||
| Amazon Nova Micro | amazon.nova-micro-v1:0 | Amazon | United States of America | 2024-12-03 | API access | 2024-12-03 | Unknown | amazon-nova-micro | |||||||
| Claude 1.3 | claude-1.3 | Anthropic | United States of America | 2023-04-18 | API access | 2023-04-18 | Unknown | An updated version of Anthropic's first generation Claude model. | claude-1-3 | ||||||
| c4ai-command-r-08-2024 | 2024-08-30 | Unverified | c4ai-command-r-08-2024 | ||||||||||||
| Command R+ | c4ai-command-r-plus-08-2024 | Cohere,Cohere for AI | Canada | 2024-08-30 | Open weights (non-commercial) | 2024-04-04 | Confident | Command R+ | command-r | ||||||
| Falcon 2 11B | falcon-11b | Falcon 2-11B | Technology Innovation Institute | United Arab Emirates | 2024-05-09 | Open weights (restricted use) | 2024-05-09 | 3.6e+23 | Confident | falcon-2-11b | 107.81 | 96.68 | 115.37 | ||
| Gemini 1.5 Flash (May 2024) | gemini-1.5-flash-0514 | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-05-14 | API access | 2024-05-10 | Unknown | gemini-1-5-flash-may-2024 | 122.38 | 115.04 | 125.34 | ||||
| Gemini 2.0 Flash | gemini-2.0-flash-exp | Google DeepMind,Google | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-12-11 | API access | 2024-12-11 | Unknown | A December 2024 experimental version of Gemini 2.0 Flash, a small langauge model from Google DeepMind. | gemini-2-0-flash | 134.87 | 131.18 | 136.87 | |||
| GPT-4 Turbo (Nov 2023) | gpt-4-turbo | OpenAI | United States of America | 2023-11-06 | API access | 2023-11-06 | Unknown | An updated version of GPT-4 and OpenAI's then-new flagship language model. | gpt-4-turbo-nov-2023 | ||||||
| Llama 3.2 11B | Llama-3.2-11B-Vision-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 5.8e+23 | Confident | An 11 billion-parameter instruction-tuned vision language model in Meta's Llama 3.2 series. | llama-3-2-11b | |||||
| mistral-small-2402 | 2024-02-26 | Unverified | mistral-small-2402 | ||||||||||||
| Mixtral 8x22B | Mixtral-8x22B-v0.1 | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | Speculative | mixtral-8x22b | 120.88 | 106.97 | 124.40 | |||
| Qwen1.5-7B | Qwen1.5-7B | Alibaba | China | 2024-02-04 | Open weights (unrestricted) | 2024-02-04 | 1.7e+23 | Confident | A 7 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-7B | qwen1-5-7b | ||||
| Qwen1.5-14B | qwen1.5-14B | Alibaba | China | 2024-02-04 | Open weights (unrestricted) | 2024-02-04 | 3.4e+23 | Confident | A 14 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-14B | qwen1-5-14b | ||||
| Qwen1.5-32B | qwen1.5-32B | Alibaba | China | 2024-02-04 | Open weights (restricted use) | 2024-02-05 | Confident | A 32 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-32B | qwen1-5-32b | |||||
| Qwen2.5-14B | qwen2.5-14b-instruct | Alibaba | China | 2025-02-26 | Open weights (unrestricted) | 2024-09-19 | 1.6e+24 | Confident | An instruction-tuned version of Qwen 2.5 14b. | Qwen2.5-14B (instruct) | qwen2-5-14b | ||||
| Qwen2.5-7B | qwen2.5-7b-instruct | Alibaba | China | 2025-02-26 | Open weights (unrestricted) | 2024-09-19 | 8.2e+23 | Confident | An instruction-tuned version of Qwen 2.5 7b. | Qwen2.5-7B | qwen2-5-7b | ||||
| Yi-Large | Yi-large | 01.AI | China | 2024-05-13 | API access | 2024-05-13 | 1.8e+24 | Speculative | Yi-Large | yi-large | |||||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_2K | Claude Sonnet 4.5 (2k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (2k thinking) | claude-sonnet-4-5 | 146.42 | 143.54 | 148.24 | ||
| GPT-5 | gpt-5-2025-08-07_low | GPT-5 (low) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using low reasoning effort. | GPT-5 (low) | gpt-5 | 150.00 | ||
| GPT-5 | gpt-5-2025-08-07_minimal | GPT-5 (minimal) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | 6.6e+25 | Speculative | OpenAI's latest flagship frontier model, benchmarked using minimal reasoning effort. | GPT-5 (minimal) | gpt-5 | 150.00 | ||
| Claude Opus 4 | claude-opus-4-20250514_2K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 2,000 reasoning tokens. | Claude Opus 4 (2k thinking) | claude-opus-4 | 143.06 | 140.01 | 145.22 | |
| Claude Sonnet 4 | claude-sonnet-4-20250514_2K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized Claude 4 series model, evaluated at up to 2,000 reasoning tokens. | Claude Sonnet 4 (2k thinking) | claude-sonnet-4 | 142.30 | 139.81 | 144.40 | ||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_2K | Claude 3.7 Sonnet (2k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 2,000 reasoning tokens allowed. | Claude 3.7 Sonnet (2k thinking) | claude-3-7-sonnet | 141.64 | 138.97 | 143.73 |
| Claude Opus 4.5 | claude-opus-4-5-20251101_64K | Claude Opus 4.5 (64k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (64k thinking) | claude-opus-4-5 | 149.93 | 146.69 | 152.48 | ||
| o3-pro | o3-pro-2025-06-10_high | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | Unknown | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with high reasoning effort. | o3-pro | 148.75 | 145.61 | 153.95 | ||||
| o3-pro | o3-pro-2025-06-10_medium | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | Unknown | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with medium reasoning effort. | o3-pro | 148.75 | 145.61 | 153.95 | ||||
| o3-pro | o3-pro-2025-06-10_low | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | Unknown | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with low reasoning effort. | o3-pro | 148.75 | 145.61 | 153.95 | ||||
| Gemini 2.5 Flash (May 2025) | gemini-2.5-flash-preview-05-20_16K | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-05-20 | API access | 2025-05-20 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 16,000 reasoning tokens. | gemini-2-5-flash-may-2025 | 141.92 | 139.59 | 143.67 | |||
| Gemini 2.5 Flash (Apr 2025) | gemini-2.5-flash-preview-04-17 (24K thinking) | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-04-17 | API access | 2025-04-17 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 24,000 reasoning tokens. | gemini-2-5-flash-apr-2025 | 139.67 | 135.33 | 143.12 | |||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro-preview-06-05_1K | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-06-05 | API access | 2025-06-05 | Unknown | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model, evaluated with up to 1,000 reasoning tokens. | Gemini 2.5 Pro Preview (Jun 2025, 1k thinking) | gemini-2-5-pro-jun-2025 | 146.18 | 144.16 | 149.21 | |
| Claude Opus 4 | claude-opus-4-20250514_8K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 8,000 reasoning tokens. | Claude Opus 4 (8k thinking) | claude-opus-4 | 143.06 | 140.01 | 145.22 | |
| Claude Sonnet 4 | claude-sonnet-4-20250514_8K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Claude Sonnet 4 (8k thinking) | claude-sonnet-4 | 142.30 | 139.81 | 144.40 | |||
| Claude Sonnet 4 | claude-sonnet-4-20250514_1K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized Claude 4 series model, evaluated at up to 1,000 reasoning tokens. | Claude Sonnet 4 (1k thinking) | claude-sonnet-4 | 142.30 | 139.81 | 144.40 | ||
| codex-mini-2025-05-16 | 2025-05-16 | Unverified | codex-mini-2025-05-16 | ||||||||||||
| o1 | o1-2024-12-17_low | o1 (low) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | Unknown | OpenAI's first reasoning model, evaluated with at the low reasoning tokens level. | o1 | 142.28 | 140.02 | 143.92 | ||
| Claude Opus 4 | claude-opus-4-20250514_1K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 1,000 reasoning tokens. | Claude Opus 4 (1k thinking) | claude-opus-4 | 143.06 | 140.01 | 145.22 | |
| Gemini 2.5 Flash (May 2025) | gemini-2.5-flash-preview-05-20_8K | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-05-20 | API access | 2025-05-20 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 8,000 reasoning tokens. | gemini-2-5-flash-may-2025 | 141.92 | 139.59 | 143.67 | |||
| o1 | o1-pro-2025-03-19_low | OpenAI | United States of America | 2025-03-19 | API access | 2024-12-05 | Unknown | An An enhanced version of OpenAI's first reasoning model,o1, evaluated with low reasoning effort. | o1 | 142.28 | 140.02 | 143.92 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_8K | Claude 3.7 Sonnet (8k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 9,000 reasoning tokens allowed. | Claude 3.7 Sonnet (8k thinking) | claude-3-7-sonnet | 141.64 | 138.97 | 143.73 |
| Gemini 2.5 Flash (May 2025) | gemini-2.5-flash-preview-05-20_1K | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-05-20 | API access | 2025-05-20 | Unknown | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 1,000 reasoning tokens. | gemini-2-5-flash-may-2025 | 141.92 | 139.59 | 143.67 | |||
| o3-mini | o3-mini-2025-01-31_low | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | Unknown | A smaller version of OpenAI's o3 reasoning model, evaluated with low reasoning effort. | o3-mini | 141.00 | 138.02 | 142.82 | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_1K | Claude 3.7 Sonnet (1k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 1,000 reasoning tokens allowed. | Claude 3.7 Sonnet (1k thinking) | claude-3-7-sonnet | 141.64 | 138.97 | 143.73 |
| Grok 3 | grok-3 | xAI | United States of America | 2025-04-09 | API access | 2025-02-17 | 3.5e+26 | Likely | XAI's third generation flagship model. | Grok 3 | grok-3 | 138.74 | 135.43 | 140.06 | |
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro-preview-06-05_32K | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | API access | 2025-06-05 | Unknown | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model, evaluated with up to 32,000 reasoning tokens. | Gemini 2.5 Pro Preview (Jun 2025, 32k thinking) | gemini-2-5-pro-jun-2025 | 146.18 | 144.16 | 149.21 | ||
| DeepSeek-V3.2-Exp | DeepSeek-V3.2-Exp_thinking | DeepSeek | China | 2025-09-29 | Open weights (unrestricted) | 2025-09-29 | 3.8e+24 | Confident | deepseek-v3-2-exp | 144.66 | 141.39 | 148.55 | |||
| Claude Opus 4 | claude-opus-4-20250514_32K | Claude Opus 4 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 32,000 reasoning tokens. | Claude Opus 4 (32k thinking) | claude-opus-4 | 143.06 | 140.01 | 145.22 | |
| Qwen3-235B-A22B | Qwen3-235B-A22B | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Likely | Qwen3 (Apr 2025) | qwen3-235b-a22b | 139.35 | 134.19 | 141.30 | ||
| Gemini 2.5 Flash (May 2025) | gemini-2.5-flash-preview-05-20_23K | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-05-20 | API access | 2025-05-20 | Unknown | A May 2025 preview version of Gemini 2.5 Flash, a small model in Google DeepMind's Gemini 2.5 series, evaluated with up to 23,000 reasoning tokens. | gemini-2-5-flash-may-2025 | 141.92 | 139.59 | 143.67 | |||
| GPT-4o (Mar 2025) | chatgpt-4o-03-27-2025 | OpenAI | United States of America | 2025-03-27 | API access | 2025-03-27 | Speculative | A version of GPT-4o that was behind the ChatGPT interface, released in March 2025. | ChatGPT-4o (Mar 2025) | gpt-4o-mar-2025 | |||||
| gpt-oss-120b | gpt-oss-120b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | The larger one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-120b | 139.64 | 135.24 | 143.10 | ||
| Qwen3-32B | Qwen3-32B | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 7.1e+24 | Confident | The 32 billion-parameter model in the Qwen 3 series. | Qwen3 32B | qwen3-32b | ||||
| Gemini 2.0 Flash | gemini-exp-1206 | Google DeepMind,Google | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-12-06 | API access | 2024-12-11 | Unknown | A December 2024 experimental version of Gemini 2.0 Flash, a small langauge model from Google DeepMind. | gemini-2-0-flash | 134.87 | 131.18 | 136.87 | |||
| GPT-4o (Jan 2025) | chatgpt-4o-01-29-2025 | OpenAI | United States of America | 2025-01-29 | API access | 2025-01-29 | Speculative | A version of GPT-4o that was behind the ChatGPT interface, released in January 2025. | ChatGPT-4o (Jan 2025) | gpt-4o-jan-2025 | |||||
| QwQ-32B | QwQ-32B | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B | qwq-32b | |||||
| DeepSeek-V2.5 | DeepSeek-V2.5 | DeepSeek | China | 2024-09-06 | Open weights (restricted use) | 2024-09-06 | 1.8e+24 | Confident | DeepSeek-V2.5 (Sept 2024) | deepseek-v2-5 | |||||
| Yi-Lightning | yi-lightning | 01.AI | China | 2024-12-02 | API access | 2024-10-18 | 1.5e+24 | Confident | Yi-Lightning | yi-lightning | |||||
| Cohere Command A | c4ai-command-a-03-2025 | Cohere | Canada | 2025-03-13 | Open weights (non-commercial) | 2025-03-13 | Confident | Command A | cohere-command-a | ||||||
| Codestral | codestral-2501 | Mistral AI | France | 2025-01-13 | Open weights (non-commercial) | 2024-05-29 | Confident | A January 2025 version of a 2024 coding-optimized model from Mistral. | codestral | ||||||
| openhands-lm-32b-v0.1 | 2024-03-26 | Unverified | openhands-lm-32b-v0-1 | ||||||||||||
| GPT-3 Medium | ada | OpenAI | United States of America | 2020-06-22 | 6.4e+20 | Confident | gpt-3-medium | ||||||||
| GPT-3 XL | babbage | OpenAI | United States of America | 2020-06-22 | 2.4e+21 | Confident | gpt-3-xl | ||||||||
| BLOOM-176B | bloom | Hugging Face,BigScience | United States of America,France | 2022-07-06 | Open weights (restricted use) | 2022-07-11 | 3.7e+23 | Confident | bloom-176b | ||||||
| GPT-3 6.7B | curie | OpenAI | United States of America | 2020-06-22 | 1.2e+22 | Confident | gpt-3-6-7b | ||||||||
| GPT-3 175B (davinci) | davinci | OpenAI | United States of America | API access | 2020-05-28 | 3.1e+23 | Confident | GPT-3 175B (davinci) | gpt-3-175b-davinci | ||||||
| GPT-3.5 | text-davinci-003 | OpenAI | United States of America | 2022-11-28 | API access | 2022-11-28 | 2.6e+24 | Speculative | gpt-3-5 | ||||||
| GPT-4 (Mar 2023) | gpt-4-32k-0314 | OpenAI | United States of America | 2023-03-14 | API access | 2023-03-15 | 2.1e+25 | Likely | gpt-4-mar-2023 | 125.76 | 118.56 | 130.42 | |||
| OPT-66B | opt-66b | Meta AI | United States of America | 2022-05-03 | Open weights (non-commercial) | 2022-06-21 | 1.1e+23 | Confident | Facebook's 66 billion-parameter model in the OPT series. | opt-66b | |||||
| OPT-175B | opt-175b | Meta AI | United States of America | 2022-05-02 | Open weights (non-commercial) | 2022-05-02 | 4.3e+23 | Confident | Facebook's 66 billion-parameter model in the OPT series. | opt-175b | |||||
| InstructGPT 350M | text-ada-001 | OpenAI | United States of America | 2022-01-27 | Confident | instructgpt-350m | |||||||||
| InstructGPT 1.3B | text-babbage-001 | OpenAI | United States of America | API access | 2022-01-27 | Confident | instructgpt-1-3b | ||||||||
| InstructGPT 6B | text-curie-001 | OpenAI | United States of America | API access | 2022-01-27 | Confident | instructgpt-6b | ||||||||
| ml-elephant | Zoo Ml-Ephant | Unverified | ml-elephant | ||||||||||||
| GPT‑5-Codex | gpt-5-codex | GPT-5-codex | OpenAI | United States of America | 2025-09-15 | API access | 2025-09-15 | Unknown | GPT-5-codex | gpt5-codex | |||||
| Gemini 2.5 Pro (Jun 2025) | gemini-2.5-pro_16K | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-06-17 | API access | 2025-06-05 | Unknown | Google’s general-purpose frontier reasoning model, evaluated with up to 16,000 reasoning tokens. | Gemini 2.5 Pro (16k thinking) | gemini-2-5-pro-jun-2025 | 146.18 | 144.16 | 149.21 | ||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_16K | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Unknown | Claude Haiku 4.5 (16k thinking) | claude-haiku-4-5 | 140.45 | 137.19 | 142.48 | |||
| Grok 4 Fast | grok-4-fast | xAI | United States of America | 2025-09-19 | API access | 2025-09-19 | Unknown | Grok 4 Fast | grok-4-fast | ||||||
| Kimi K2 Thinking | kimi-k2-thinking | kimi-k2-thinking (official) | Moonshot | China | 2025-11-06 | Open weights (restricted use) | 2025-11-06 | 4.2e+24 | Likely | kimi-k2-thinking | 145.05 | 142.80 | 146.99 | ||
| Grok-3 mini | grok-3-mini_high | Grok 3 mini | xAI | United States of America | 2025-06-24 | API access | 2025-02-19 | Unknown | Grok 3 mini (high) | grok-3-mini | 140.70 | 137.75 | 142.61 | ||
| Gemini 2.5 Flash (Sep 2025) | gemini-2.5-flash-preview-09-2025_16k | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-09-25 | API access | 2025-09-25 | Unknown | Gemini 2.5 Flash-Lite (Sep 2025, 16k thinking) | gemini-2-5-flash-sep-2025 | 142.11 | 137.83 | 144.32 | |||
| gpt-oss-120b | gpt-oss-120b_medium | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | The larger one of OpenAI's two latest open source language models, evaluated with medium reasoning effort. | gpt-oss-120b | 139.64 | 135.24 | 143.10 | ||
| gpt-oss-20b | gpt-oss-20b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | Confident | The smaller one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-20b | |||||
| GLM-4.5 | glm-4.5_thinking | GLM-4.5 Thinking | Z.ai (Zhipu AI),Tsinghua University | China | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.4e+24 | Confident | glm-4-5 | |||||
| GPT-5 | gpt-5-chat | OpenAI | United States of America | API access | 2025-08-07 | 6.6e+25 | Speculative | gpt-5 | 150.00 | ||||||
| Qwen3-235B-A22B-Instruct (Jul 2025) | Qwen3-235B-A22B-Instruct-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | A 22 billion active (235 billion total)-parameter mixture-of-experts instruction-tuned model in the Qwen 3 series. | Qwen3 Non-thinking (Jul 2025) | qwen3-235b-a22b-instruct-jul-2025 | ||||
| DeepSeek-V3.1 | DeepSeek-V3.1_thinking | DeepSeek | China | 2025-08-21 | Open weights (unrestricted) | 2025-08-21 | 3.6e+24 | Confident | DeepSeek-V3.1 (thinking) | deepseek-v3-1 | 138.53 | 135.09 | 142.99 | ||
| gpt-oss-20b | gpt-oss-20b_medium | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | Confident | The smaller one of OpenAI's two latest open source language models, evaluated with medium reasoning effort. | gpt-oss-20b | |||||
| Kimi K2 | Kimi-K2-Instruct-0905 | Moonshot | China | 2024-09-05 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | kimi-k2 | 140.37 | 137.40 | 142.49 | |||
| Gemini 2.5 Flash-Lite (Jun 2025) | gemini-2.5-flash-lite-preview-06-17_16K | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-06-17 | API access | 2025-06-15 | Unknown | A June 2025 preview of Gemini 2.5 Flash Lite, evaluated with up to 16,000 reasoning tokens. | Gemini 2.5 Flash-Lite (Jun 2025, 16k thinking) | gemini-2-5-flash-lite-jun-2025 | |||||
| grok-code-fast-1 | 2025-08-28 | Unverified | grok-code-fast-1 | ||||||||||||
| mistral-medium-2508 | 2025-08-12 | Unverified | mistral-medium-2508 | ||||||||||||
| Qwen3-30B-A3B | Qwen3-30B-A3B | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 6.5e+23 | Likely | Qwen3-30B-A3B | qwen3-30b-a3b | |||||
| Llama 3-8B | Meta-Llama-3-8B | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.2e+23 | Confident | Meta's 8 billion-parameter model in the Llama 3 series. | llama-3-8b | 116.00 | 105.44 | 119.81 | ||
| Llama 3-70B | Meta-Llama-3-70B | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.9e+24 | Confident | Meta's 70 billion-parameter model in the Llama 3 series. | llama-3-70b | 119.30 | 87.91 | 125.18 | ||
| Reka Flash 3 | reka-flash-3 | Reka AI | United States of America | 2025-03-10 | Open weights (unrestricted) | 2025-03-10 | Confident | Reka Flash 3 | reka-flash-3 | ||||||
| DeepSeek-R1-Distill-Qwen-32B | DeepSeek-R1-Distill-Qwen-32B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | Speculative | A model based on Qwen 32B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | deepseek-r1-distill-qwen-32b | ||||||
| Mistral NeMo | Mistral-Nemo-Instruct-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | Confident | mistral-nemo | 117.94 | 109.61 | 121.82 | ||||
| Llama 3.2 3B | Llama-3.2-3B-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 1.7e+23 | Confident | A 3 billion-parameter instruction-tuned model in Meta's Llama 3.2 series. | llama-3-2-3b | |||||
| Llama 3.2 1B | Llama-3.2-1B-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 6.6e+22 | Confident | A 1 billion-parameter instruction-tuned model in Meta's Llama 3.2 series. | llama-3-2-1b | |||||
| Claude Opus 4 | claude-opus-4-20250514_12K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 12,000 reasoning tokens. | Claude Opus 4 (12k thinking) | claude-opus-4 | 143.06 | 140.01 | 145.22 | ||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_12K | Claude Sonnet 4.5 (12k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Unknown | Claude Sonnet 4.5 (12k thinking) | claude-sonnet-4-5 | 146.42 | 143.54 | 148.24 | ||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_12K | Claude 3.7 Sonnet (12k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 12,000 reasoning tokens allowed. | Claude 3.7 Sonnet (12k thinking) | claude-3-7-sonnet | 141.64 | 138.97 | 143.73 |
| Claude Sonnet 4 | claude-sonnet-4-20250514_12K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Unknown | Anthropic's mid-sized Claude 4 series model, evaluated at up to 12,000 reasoning tokens. | Claude Sonnet 4 (12k thinking) | claude-sonnet-4 | 142.30 | 139.81 | 144.40 | ||
| GLM-4.6 | glm-4.6 | Z.ai (Zhipu AI),Tsinghua University | China | 2025-09-30 | Open weights (unrestricted) | 2025-09-30 | 4.4e+24 | Likely | glm-4-6 | ||||||
| GPT-5.1-codex | gpt-5.1-codex | GPT-5.1-codex | OpenAI | United States of America | 2025-11-12 | API access | 2025-11-12 | Unverified | GPT-5.1-codex | gpt-5-1-codex | |||||
| DeepSeek-V3.2 | deepseek/deepseek-v3.2 | DeepSeek | China | 2025-12-01 | 2025-12-01 | Unverified | deepseek-v3-2 | 144.60 | 140.97 | 147.18 | |||||
| GPT-5.1-codex-mini | gpt-5.1-codex-mini | GPT-5.1-codex mini | OpenAI | United States of America | 2025-11-12 | API access | 2025-11-12 | Unverified | GPT-5.1-codex mini | gpt-5-1-codex-mini | |||||
| Grok 4.1 Fast | grok-4-1-fast-reasoning | xAI | United States of America | API access | 2025-11-19 | Unverified | grok-4-1-fast | ||||||||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_15K | Claude 3.7 Sonnet (15k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | Likely | A previous flagship reasoning model from Anthropic, benchmarked with up to 15,000 reasoning tokens allowed. | Claude 3.7 Sonnet (15k thinking) | claude-3-7-sonnet | 141.64 | 138.97 | 143.73 |
| Pixtral 12B | Pixtral-12B-2409 | Mistral AI | France | 2024-09-17 | Open weights (unrestricted) | 2024-09-17 | Confident | pixtral-12b | |||||||
| GPT-5.1-Codex-Max | gpt-5.1-codex-max | GPT-5.1-Codex-Max | OpenAI | United States of America | 2025-11-19 | API access | 2025-11-19 | Unverified | GPT-5.1-Codex-Max | gpt-5-1-codex-max | |||||
| Claude Opus 4.5 | claude-opus-4-5-20251101_128K | Claude Opus 4.5 (128k thinking) | Anthropic | United States of America | 2025-11-24 | API access | 2025-11-24 | Unknown | Claude Opus 4.5 (64k thinking) | claude-opus-4-5 | 149.93 | 146.69 | 152.48 | ||
| gpt-oss-20b | gpt-oss-20b | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | Confident | The smaller one of OpenAI's two latest open source language models. | gpt-oss-20b | |||||
| MPT-30B | mpt-30b-instruct | MosaicML | United States of America | 2023-06-22 | Open weights (unrestricted) | 2023-06-22 | 1.9e+23 | Confident | mpt-30b | 97.24 | 82.73 | 104.72 | |||
| Vicuna-13B-v1.3 | vicuna-13b-v1.3 | Large Model Systems Organization,University of California (UC) Berkeley | United States of America | 2023-06-18 | Open weights (restricted use) | 2023-06-22 | Confident | Vicuna-13B-v1.3 | vicuna-13b-v1-3 | ||||||
| T5-Small | Unverified | t5-small | |||||||||||||
| T5-Base | Unverified | t5-base | |||||||||||||
| Switch-Base | Unverified | switch-base | |||||||||||||
| Switch-Large | Unverified | switch-large | |||||||||||||
| GPT-3.5 Turbo | gpt-3.5-turbo-instruct | OpenAI | United States of America | 2023-09-18 | API access | 2022-11-30 | Speculative | An instruction-tuned version of OpenAI's GPT-3.5 turbo. | gpt-3-5-turbo | 116.62 | 105.78 | 121.73 | |||
| GPT-3 175B (davinci) | davinci-002 | OpenAI | United States of America | API access | 2020-05-28 | 3.1e+23 | Confident | GPT-3 175B (davinci-002) | gpt-3-175b-davinci | ||||||
| GPT-3.5 | code-davinci-002 | OpenAI | United States of America | API access | 2022-11-28 | 2.6e+24 | Speculative | gpt-3-5 | |||||||
| Falcon-40B | falcon-40b-instruct | Technology Innovation Institute | United Arab Emirates | 2023-05-25 | Open weights (unrestricted) | 2023-03-15 | 2.4e+23 | Confident | falcon-40b | 101.71 | 89.28 | 109.80 | |||
| DeepSeek-Coder-V2-Lite-Instruct | 2024-06-13 | Unverified | deepseek-coder-v2-lite-instruct | ||||||||||||
| DeepSeek-Coder-V2 236B | DeepSeek-Coder-V2-Instruct | DeepSeek | China | 2024-06-17 | Open weights (restricted use) | 2024-06-17 | 1.3e+24 | Confident | An instruction-tuned version of DeepSeek Coder V2, a 236 billion-parameter coding-optimized model from DeepSeek. | DeepSeek-Coder-V2 (instruct) | deepseek-coder-v2-236b | ||||
| Qwen2.5-Coder-3B-Instruct | 2024-11-06 | Unverified | qwen2-5-coder-3b-instruct | ||||||||||||
| Qwen2.5-Coder (7B) | Qwen2.5-Coder-7B-Instruct | Alibaba | China | 2024-09-17 | Open weights (unrestricted) | 2024-09-18 | 2.5e+23 | Confident | An instruction-tuned and coding-optimized 7 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-Coder-7B (instruct) | qwen2-5-coder-7b | 111.89 | 95.03 | 119.46 | |
| Qwen2.5-Coder-14B-Instruct | 2024-11-06 | Unverified | qwen2-5-coder-14b-instruct | ||||||||||||
| QwQ-32B | QwQ-32B (16K thinking) | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B (16k thinking) | qwq-32b | |||||
| Qwen2.5-Max | qwen2.5-max | Alibaba | China | 2025-01-28 | API access | 2025-01-28 | Unknown | The largest model in the Qwen 2.5 series. | Qwen2.5-Max | qwen2-5-max | 132.98 | 127.68 | 136.39 | ||
| Kimi K2 | moonshotai/kimi-k2-0905 | Kimi K2 0905 (Novita) | Moonshot | China | 2025-09-05 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | A July 2025 preview version of Kimi K2 turbo, an updated version of Kimi K2. | kimi-k2 | 140.37 | 137.40 | 142.49 | |
| gemini-robotics-er-1.5-preview | 2025-09-26 | Unverified | gemini-robotics-er-1-5-preview | ||||||||||||
| Gemini 2.5 Flash-Lite (Sep 2025) | gemini-2.5-flash-lite-preview-09-2025 | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-09-25 | API access | 2025-09-25 | Unknown | Gemini 2.5 Flash-Lite (Sep 2025) | gemini-2-5-flash-lite-sep-2025 | ||||||
| computer-use-preview-2025-03-11 | 2025-03-11 | Unverified | computer-use-preview-2025-03-11 | ||||||||||||
| Gemini 2.0 Flash-Lite (Feb 2024) | gemini-2.0-flash-lite | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-02-05 | API access | 2024-02-05 | Unknown | The smallest model in Google DeepMind's Gemini 2.0 series. | gemini-2-0-flash-lite-feb-2024 | ||||||
| Gemini 2.0 Flash-Lite (Feb 2024) | gemini-2.0-flash-lite-preview-02-05 | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-02-05 | API access | 2024-02-05 | Unknown | A February 2025 preview version of the smallest model in Google DeepMind's Gemini 2.0 series. | gemini-2-0-flash-lite-feb-2024 | ||||||
| Dracarys2-72B-Instruct | 2024-09-30 | Unverified | dracarys2-72b-instruct | ||||||||||||
| learnlm-1.5-pro-experimental | Unverified | learnlm-1-5-pro-experimental | |||||||||||||
| sonar | Perplexity Sonar | Unverified | sonar | ||||||||||||
| Dracarys2-Llama-3.1-70B-Instruct | 2024-08-14 | Unverified | dracarys2-llama-3-1-70b-instruct | ||||||||||||
| QwQ-32B | QwQ-32B-Preview | Alibaba | China | 2024-11-28 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B Preview | qwq-32b | |||||
| OLMo 2 Furious 13B | OLMo-2-1124-13B-Instruct | Allen Institute for AI,University of Washington,New York University (NYU) | United States of America | 2024-12-31 | Open weights (unrestricted) | 2024-12-31 | 4.6e+23 | Confident | olmo-2-furious-13b | ||||||
| BLIP-2 (Q-Former) | blip2-opt-2.7b | Salesforce Research | United States of America | 2023-02-06 | Open weights (unrestricted) | 2023-01-30 | 1.2e+21 | Confident | blip-2-q-former | ||||||
| Phi-3.5-vision-instruct | 2024-08-16 | Unverified | phi-3-5-vision-instruct | ||||||||||||
| MM1-3B-Chat | 2024-03-14 | Unverified | mm1-3b-chat | ||||||||||||
| MM1-7B-Chat | 2024-03-14 | Unverified | mm1-7b-chat | ||||||||||||
| llava-v1.6-vicuna-7b | 2024-01-31 | Unverified | llava-v1-6-vicuna-7b | ||||||||||||
| llama3-llava-next-8b | 2024-04-20 | Unverified | llama3-llava-next-8b | ||||||||||||
| Gemini 1.0 Pro Vision | gemini-1.0-pro-vision | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2024-01-04 | API access | 2024-02-15 | Unknown | A vision language model from Google DeepMind. | gemini-1-0-pro-vision | ||||||
| falcon-11B-vlm | 2024-05-21 | Unverified | falcon-11b-vlm | ||||||||||||
| llava-v1.6-vicuna-13b | 2024-01-31 | Unverified | llava-v1-6-vicuna-13b | ||||||||||||
| llava-v1.6-mistral-7b | 2024-01-31 | Unverified | llava-v1-6-mistral-7b | ||||||||||||
| instructblip-vicuna-7b | 2023-05-22 | Unverified | instructblip-vicuna-7b | ||||||||||||
| instructblip-vicuna-13b | 2023-12-25 | Unverified | instructblip-vicuna-13b | ||||||||||||
| InternVL-Chat-ViT-6B-Vicuna-7B | 2023-12-25 | Unverified | internvl-chat-vit-6b-vicuna-7b | ||||||||||||
| InternVL-Chat-ViT-6B-Vicuna-13B | 2024-08-16 | Unverified | internvl-chat-vit-6b-vicuna-13b | ||||||||||||
| llava-v1.5-7b | 2023-10-05 | Unverified | llava-v1-5-7b | ||||||||||||
| DeepSeek-R1 (May 2025) | chutes/DeepSeek-R1-0528 | DeepSeek-R1 (May 2025) | DeepSeek | China | 2025-05-28 | Open weights (unrestricted) | 2025-05-28 | 4.0e+24 | Confident | DeepSeek-R1 (May 2025) (Chutes) | deepseek-r1-may-2025 | 141.75 | 139.03 | 143.81 | |
| chutes/GLM-4.5-FP8 | 2025-07-27 | Unverified | chutes-glm-4-5-fp8 | ||||||||||||
| gpt-oss-120b | chutes/gpt-oss-120b | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | gpt-oss-120b (Chutes) | gpt-oss-120b | 139.64 | 135.24 | 143.10 | ||
| gpt-oss-120b | chutes/gpt-oss-120b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | Confident | gpt-oss-120b (high) (Chutes) | gpt-oss-120b | 139.64 | 135.24 | 143.10 | ||
| Llama 4 Maverick | chutes/Llama-4-Maverick-17B-128E-Instruct | Meta AI | United States of America | 2025-04-06 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | Likely | Llama 4 Maverick (128 experts) (Chutes) | llama-4-maverick | 132.69 | 126.70 | 134.91 | ||
| Qwen3-8B | chutes/Qwen3-8B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 1.8e+24 | Confident | Qwen3 8B (Chutes) | qwen3-8b | |||||
| Qwen3-235B-A22B | chutes/Qwen3-235B-A22B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Likely | Qwen3 (Apr 2025) (Chutes) | qwen3-235b-a22b | 139.35 | 134.19 | 141.30 | ||
| Qwen3-235B-A22B-Thinking (Jul 2025) | chutes/Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | Qwen3 (Jul 2025) (Chutes) | qwen3-235b-a22b-thinking-jul-2025 | 145.26 | 141.93 | 148.42 | ||
| QwQ-32B | chutes/QwQ-32B | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | Speculative | QwQ 32B (Chutes) | qwq-32b | |||||
| deepinfra/Qwen3-Next-80B-A3B-Instruct | Unverified | deepinfra-qwen3-next-80b-a3b-instruct | |||||||||||||
| DeepSeek-V3.2-Exp | DeepSeek-V3.2-Exp_high | DeepSeek | China | 2025-09-29 | Open weights (unrestricted) | 2025-09-29 | 3.8e+24 | Confident | deepseek-v3-2-exp | 144.66 | 141.39 | 148.55 | |||
| Gemini 2.5 Flash-Lite (Jun 2025) | gemini-2.5-flash-lite-preview-06-17-thinking | gemini-2.5-flash-lite-preview-06-17 (Default thinking length) | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-06-17 | API access | 2025-06-15 | Unknown | A June 2025 preview of Gemini 2.5 Flash Lite. | Gemini 2.5 Flash-Lite (Jun 2025) | gemini-2-5-flash-lite-jun-2025 | ||||
| DeepSeek-V3 | chutes/DeepSeek-V3-0324 | DeepSeek-V3 (Mar 2025) | DeepSeek | China | 2025-03-24 | Open weights (restricted use) | 2024-12-24 | 3.4e+24 | Confident | DeepSeek-V3 (Mar 2025) (Chutes) | deepseek-v3 | 136.51 | 132.41 | 138.30 | |
| Gemma 3 27B | chutes/Gemma-3-27b-It | Google DeepMind | United States of America,United Kingdom of Great Britain and Northern Ireland | 2025-03-11 | Open weights (restricted use) | 2025-03-12 | 2.3e+24 | Confident | Gemma-3-27b-it (Chutes) | gemma-3-27b | 130.72 | 123.89 | 133.03 | ||
| Kimi K2 | fireworks/Kimi-K2-Instruct-0905 | Kimi K2 Instruct (0905) | Moonshot | China | 2025-09-05 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Confident | kimi-k2 | 140.37 | 137.40 | 142.49 | ||
| Grok-3 mini | grok-3-mini-beta_medium | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | Unknown | A beta release of a smaller version of XAI's third generation Grok model, evaluated with medium reasoning effort. | grok-3-mini | 140.70 | 137.75 | 142.61 | |||
| MiniMax-M1-80k | MiniMax-M1-80k | MiniMax | China | 2025-06-13 | Open weights (unrestricted) | 2025-06-13 | 4.3e+24 | Likely | minimax-m1-80k | ||||||
| nvidia-nemotron-nano-9b-v2 | 2025-08-18 | Unverified | nvidia-nemotron-nano-9b-v2 | ||||||||||||
| Qwen3-235B-A22B-Instruct (Jul 2025) | parasail-qwen3-235b-a22b-instruct-2507 | Alibaba | China | 2025-09-01 | Open weights (unrestricted) | 2025-07-25 | 4.8e+24 | Likely | Qwen3 (Parasail) | qwen3-235b-a22b-instruct-jul-2025 | |||||
| Qwen3-30B-A3B | chutes/Qwen3-30B-A3B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 6.5e+23 | Likely | Qwen3-30B-A3B (Chutes) | qwen3-30b-a3b | |||||
| Qwen3-14B | chutes/Qwen3-14B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 3.2e+24 | Confident | Qwen3 14B | qwen3-14b | |||||
| Qwen3-32B | chutes/Qwen3-32B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 7.1e+24 | Confident | Qwen3 32B (Chutes) | qwen3-32b | |||||
| chutes/Qwen3-Next-80B-A3B-Instruct | 2025-09-11 | Unverified | chutes-qwen3-next-80b-a3b-instruct | ||||||||||||
| Llama 4 Scout | chutes/Llama-4-Scout-17B-16E Instruct | Meta AI | United States of America | 2025-04-06 | Open weights (restricted use) | 2025-04-05 | 4.1e+24 | Likely | Llama 4 Maverick (16 experts) (Chutes) | llama-4-scout | 129.82 | 123.74 | 131.82 |
This gap closely resembles the broader gap between proprietary and open-weight models. This is unsurprising since nearly all leading Chinese models are open-weight, while frontier US models remain closed.
Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.
Learn more about this graph
We visualize the gap in capabilities between US and Chinese models, using the Epoch Capabilities Index (ECI). Since 2023, the gap has ranged from 4 to 14 months, with a mean gap of 7 months.
Analysis
To calculate the gap between US and Chinese models, we first find the set of models that had the highest ECI among models from their country upon release. We then drop the first of these models (LLaMA-65B for the US, and Baichuan1-7B for China), since these first models were likely not at the true frontier (ECI data starts in January 2023).
To quantify the gap on each day, we look at the ECI of the best Chinese model on that day, and then calculate how long it has been since the last time the leading US model was the same or worse than that score. We consider models to be the same performance if their scores are within 1 ECI point difference. We repeat this process for each day where values exist for both the US and China. In practice, the first point where a Chinese model surpasses GPT-4 is May 2024 (a gap of 14 months), and no Chinese model has yet surpassed the ECI of OpenAI’s o3 model, released in April 2025.



