Frontier open-weight models lag behind the most capable models by an average of 3 months in the Epoch Capabilities Index (ECI), our holistic measure of model capability. That corresponds to an average ECI gap of around 7 points, similar to the gap between o3 and GPT-5.
| Model | Model version ID | Display name | Organization | Country (of organization) | Version release date | Model accessibility | Publication date | Training compute (FLOP) | Description | Unique display name | Slug | ECI |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| o3-mini | o3-mini-2025-01-31_high | o3-mini (high) | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | A smaller version of OpenAI's o3 reasoning model, evaluated with high reasoning effort. | o3-mini-2025-01-31-high | 140.39 | ||
| Grok-2 | grok-2-1212 | xAI | United States of America | 2024-12-12 | API access | 2024-08-13 | 3.0e+25 | XAI's second generation Grok model. | grok-2-1212 | 130.29 | ||
| o3-mini | o3-mini-2025-01-31_medium | o3-mini (medium) | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | A smaller version of OpenAI's o3 reasoning model, evaluated with medium reasoning effort. | o3-mini-2025-01-31-medium | 138.51 | ||
| Mistral Large 2 | mistral-large-2411 | Mistral AI | France | 2024-11-18 | Open weights (non-commercial) | 2024-07-24 | 2.1e+25 | mistral-large-2411 | 128.50 | |||
| GPT-4o | gpt-4o-2024-11-20 | GPT-4o (Nov 2024) | OpenAI | United States of America | 2024-11-20 | API access | 2024-05-13 | The November 2024 version of GPT-4o, OpenAI's then-flagship multimodal language model. | gpt-4o-2024-11-20 | 129.09 | ||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219 | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | A previous flagship reasoning model from Anthropic. | Claude 3.7 Sonnet | claude-3-7-sonnet-20250219 | 136.16 | |
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_16K | Claude 3.7 Sonnet (16k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | A previous flagship reasoning model from Anthropic, benchmarked with up to 16,000 reasoning tokens allowed. | Claude 3.7 Sonnet (16k thinking) | claude-3-7-sonnet-20250219-16k | 138.21 |
| o1-mini | o1-mini-2024-09-12_high | o1-mini (high) | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | A smaller version of OpenAI’s first reasoning model, o1, evaluated with high reasoning effort. | o1-mini-2024-09-12-high | 136.75 | ||
| o1-mini | o1-mini-2024-09-12_medium | o1-mini (medium) | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | A smaller version of OpenAI’s first reasoning model, o1, evaluated with medium reasoning effort. | o1-mini-2024-09-12-medium | 135.31 | ||
| Claude 3.5 Sonnet | claude-3-5-sonnet-20241022 | Claude 3.5 Sonnet (Oct 2024) | Anthropic | United States of America | 2024-10-22 | API access | 2024-06-20 | 2.7e+25 | An updated version of the top model in the Claude 3.5 series, a previous flagship model series from Anthropic. | Claude 3.5 Sonnet (Oct 2024) | claude-3-5-sonnet-20241022 | 133.55 |
| Gemini 1.5 Flash | gemini-1.5-flash-002 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-09-24 | API access | 2024-05-10 | The second version of Google DeepMind's Gemini 1.5 Flash. | gemini-1-5-flash-002 | 129.78 | |||
| Claude 3.5 Haiku | claude-3-5-haiku-20241022 | Claude 3.5 Haiku (Oct 2024) | Anthropic | United States of America | 2024-10-22 | API access | 2024-10-22 | The smallest model in Anthropic’s Claude 3.5 Series. | Claude 3.5 Haiku (Oct 2024) | claude-3-5-haiku-20241022 | 127.48 | |
| Claude 3.5 Sonnet | claude-3-5-sonnet-20240620 | Claude 3.5 Sonnet (Jun 2024) | Anthropic | United States of America | 2024-06-20 | API access | 2024-06-20 | 2.7e+25 | The first version of the top model in the Claude 3.5 series, a previous flagship model series from Anthropic. | Claude 3.5 Sonnet (Jun 2024) | claude-3-5-sonnet-20240620 | 130.00 |
| GPT-4o | gpt-4o-2024-08-06 | GPT-4o (Aug 2024) | OpenAI | United States of America | 2024-08-06 | API access | 2024-05-13 | An updated version of GPT-4o, OpenAI's model that powered ChatGPT from mid-2024 to -2025. | gpt-4o-2024-08-06 | 129.06 | ||
| Gemini 2.0 Flash | gemini-2.0-flash-001 | Gemini 2.0 Flash (Feb 2025) | Google DeepMind,Google | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-02-05 | API access | 2024-12-11 | The first version of Google DeepMind's Gemini 2.0 Flash model. | gemini-2-0-flash-001 | 134.57 | ||
| o1 | o1-2024-12-17_high | o1 (high) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | OpenAI's first reasoning model, evaluated with at the high reasoning effort level. | o1-2024-12-17-high | 140.56 | ||
| DeepSeek-V3 | DeepSeek-V3 | DeepSeek | China | 2024-12-26 | Open weights (restricted use) | 2024-12-24 | 3.4e+24 | DeepSeek’s 2024 mixture-of-experts model. | DeepSeek-V3 | deepseek-v3 | 132.34 | |
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_32K | Claude 3.7 Sonnet (32k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | A previous flagship reasoning model from Anthropic, benchmarked with up to 32,000 reasoning tokens allowed. | Claude 3.7 Sonnet (32k thinking) | claude-3-7-sonnet-20250219-32k | 139.00 |
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_64K | Claude 3.7 Sonnet (64k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | A previous flagship reasoning model from Anthropic, benchmarked with up to 64,000 reasoning tokens allowed. | Claude 3.7 Sonnet (64k thinking) | claude-3-7-sonnet-20250219-64k | 139.14 |
| Qwen2.5-Max | qwen-max-2025-01-25 | Alibaba | China | 2025-01-25 | API access | 2025-01-28 | Qwen2.5-Max | qwen-max-2025-01-25 | 132.54 | |||
| Llama 4 Scout | Llama-4-Scout-17B-16E-Instruct | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 4.1e+24 | The 17 billion active (109 billion total)-parameter model in Meta's Llama 4 series. | llama-4-scout-17b-16e-instruct | 128.91 | ||
| Llama 4 Maverick | Llama-4-Maverick-17B-128E-Instruct-FP8 | Llama 4 Maverick (FP8) | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | The 17 billion active (400 billion total)-parameter model in Meta's Llama 4 series, quantized to FP8. | llama-4-maverick-17b-128e-instruct-fp8 | 133.52 | |
| Grok 3 | grok-3-beta | xAI | United States of America | 2025-04-09 | API access | 2025-02-17 | 3.5e+26 | A beta release of XAI's third generation Grok model. | Grok 3 (beta) | grok-3-beta | 138.01 | |
| Grok-3 mini | grok-3-mini-beta_high | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | A beta release of a smaller version of XAI's third generation Grok model, evaluated with high reasoning effort. | grok-3-mini-beta-high | 139.97 | |||
| Grok-3 mini | grok-3-mini-beta_low | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | A beta release of a smaller version of XAI's third generation Grok model, evaluated with low reasoning effort. | grok-3-mini-beta-low | 137.44 | |||
| GPT-4.1 | gpt-4.1-2025-04-14 | GPT-4.1 | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | A coding-optimized GPT-4 series model from OpenAI. | gpt-4-1-2025-04-14 | 136.19 | ||
| GPT-4.1 mini | gpt-4.1-mini-2025-04-14 | GPT-4.1 mini | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | A smaller version of OpenAI's GPT-4.1. | gpt-4-1-mini-2025-04-14 | 134.60 | ||
| GPT-4.1 nano | gpt-4.1-nano-2025-04-14 | GPT-4.1 nano | OpenAI | United States of America | 2025-04-14 | API access | 2025-04-14 | The smallest version of OpenAI's GPT-4.1. | gpt-4-1-nano-2025-04-14 | 129.88 | ||
| o3 | o3-2025-04-16_high | o3 (high) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | OpenAI's second generation o-series reasoning model, evaluated with high reasoning effort. | o3-2025-04-16-high | 144.16 | ||
| o4-mini | o4-mini-2025-04-16_high | o4-mini (high) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | The latest smaller size o-series reasoning model from OpenAI, evaluated with high reasoning effort. | o4-mini-2025-04-16-high | 143.43 | ||
| o3 | o3-2025-04-16_medium | o3 (medium) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | OpenAI's second generation o-series reasoning model, evaluated with medium reasoning effort. | o3-2025-04-16-medium | 144.06 | ||
| o3 | o3-2025-04-16_low | o3 (low) | OpenAI | United States of America | 2025-04-16 | API access | 2024-12-20 | OpenAI's latest full-size o-series reasoning model release, evaluated on low reasoning effort. | o3-2025-04-16-low | |||
| o4-mini | o4-mini-2025-04-16_medium | o4-mini (medium) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | The latest smaller size o-series reasoning model from OpenAI, evaluated with medium reasoning effort. | o4-mini-2025-04-16-medium | 141.83 | ||
| o4-mini | o4-mini-2025-04-16_low | o4-mini (low) | OpenAI | United States of America | 2025-04-16 | API access | 2025-04-16 | The latest smaller size o-series reasoning model from OpenAI, evaluated with low reasoning effort. | o4-mini-2025-04-16-low | |||
| Mistral Medium 3 | mistral-medium-2505 | Mistral AI | France | 2025-05-07 | API access | 2025-05-07 | A 2025 language model from Mistral. | mistral-medium-2505 | 134.23 | |||
| Qwen Plus | qwen-plus-2025-04-28 | Alibaba | China | 2025-04-28 | API access | 2024-02-06 | Qwen Plus (Apr 2025) | qwen-plus-2025-04-28 | ||||
| Gemini 2.5 Pro | gemini-2.5-pro-preview-06-05 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-06-05 | API access | 2025-03-25 | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model. | Gemini 2.5 Pro Preview (Jun 2025) | gemini-2-5-pro-preview-06-05 | 146.12 | |
| Gemini 2.5 Pro | gemini-2.5-pro | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-06-17 | API access | 2025-03-25 | Google’s general-purpose frontier reasoning model. | Gemini 2.5 Pro | gemini-2-5-pro | |||
| Claude Opus 4 | claude-opus-4-20250514 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1). | Claude Opus 4 | claude-opus-4-20250514 | 138.27 | ||
| Claude Sonnet 4 | claude-sonnet-4-20250514 | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's mid-sized version of Claude 4 series of reasoning models. | Claude Sonnet 4 | claude-sonnet-4-20250514 | 136.22 | ||
| Grok 4 | grok-4-0709 | xAI | United States of America | 2025-07-09 | API access | 2025-07-09 | 5.0e+26 | Grok 4 | grok-4-0709 | 145.77 | ||
| Claude Opus 4.1 | claude-opus-4-1-20250805_27K | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Anthropic’s largest flagship model, benchmarked with up to 27,000 reasoning tokens. | Claude Opus 4.1 (27k thinking) | claude-opus-4-1-20250805-27k | 140.53 | ||
| Claude Opus 4.1 | claude-opus-4-1-20250805 | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Anthropic’s largest flagship model. | Claude Opus 4.1 | claude-opus-4-1-20250805 | 138.48 | ||
| Claude Opus 4 | claude-opus-4-20250514_27K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 27,000 reasoning tokens. | Claude Opus 4 (27k thinking) | claude-opus-4-20250514-27k | |||
| Kimi K2 | Kimi-K2-Instruct | Moonshot | China | 2025-07-12 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | Moonshot AI’s 1-trillion parameter scale model. | kimi-k2-instruct | 134.84 | ||
| GPT-5 | gpt-5-2025-08-07_medium | GPT-5 (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 (medium) | gpt-5-2025-08-07-medium | 150.00 | |
| GPT-5 | gpt-5-2025-08-07_high | GPT-5 (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 (high) | gpt-5-2025-08-07-high | 148.45 | |
| GPT-5 nano | gpt-5-nano-2025-08-07_medium | GPT-5 nano (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | The smallest version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 nano (medium) | gpt-5-nano-2025-08-07-medium | 139.14 | |
| GPT-5 mini | gpt-5-mini-2025-08-07_high | GPT-5 mini (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | A smaller version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 mini (high) | gpt-5-mini-2025-08-07-high | 142.30 | |
| GPT-5 nano | gpt-5-nano-2025-08-07_high | GPT-5 nano (high) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | The smallest version of OpenAI's latest flagship frontier model, benchmarked using high reasoning effort. | GPT-5 nano (high) | gpt-5-nano-2025-08-07-high | 140.10 | |
| GPT-5 mini | gpt-5-mini-2025-08-07_medium | GPT-5 mini (medium) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | A smaller version of OpenAI's latest flagship frontier model, benchmarked using medium reasoning effort. | GPT-5 mini (medium) | gpt-5-mini-2025-08-07-medium | 141.84 | |
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929 | Claude Sonnet 4.5 (no thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Anthropic's latest mid-size model, part of the Claude 4.5 series. | Claude Sonnet 4.5 (no thinking) | claude-sonnet-4-5-20250929 | 137.97 | |
| Gemini 2.5 Deep Think | gemini-2.5-deep-think-2025-08-01-webapp | Google,Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-08-01 | Hosted access (no API) | 2025-08-01 | gemini-2-5-deep-think-2025-08-01-webapp | |||||
| Claude Haiku 4.5 | claude-haiku-4-5-20251001 | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | claude-haiku-4-5-20251001 | |||||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_32K | Claude Sonnet 4.5 (32k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Anthropic's latest mid-size model, part of the Claude 4.5 series, evaluated with up to 32,000 reasoning tokens. | Claude Sonnet 4.5 (32k thinking) | claude-sonnet-4-5-20250929-32k | 142.39 | |
| Claude Haiku 4.5 | claude-haiku-4-5-20251001_32K | Anthropic | United States of America | 2025-10-15 | API access | 2025-10-15 | Claude Haiku 4.5 (32k thinking) | claude-haiku-4-5-20251001-32k | ||||
| Gemini 2.0 Pro | gemini-2.0-pro-exp-02-05 | Gemini 2.0 Pro Exp (Feb 2025) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-02-05 | Hosted access (no API) | 2024-12-11 | A February 2025 experimental version of Google DeepMind's previous flagship model, Gemini 2.0 Pro. | gemini-2-0-pro-exp-02-05 | 134.45 | ||
| DeepSeek-V3 | DeepSeek-V3-0324 | DeepSeek-V3 (Mar 2025) | DeepSeek | China | 2025-03-24 | Open weights (restricted use) | 2024-12-24 | 3.4e+24 | An updated version of the DeepSeek V3 model. | DeepSeek-V3 (Mar 2025) | deepseek-v3-0324 | 136.16 |
| Qwen3-235B-A22B | qwen3-235b-a22b | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Qwen3 (Apr 2025) | qwen3-235b-a22b | 135.50 | ||
| Magistral Small 1.1 | magistral-small-2506 | Mistral AI | France | 2025-06-10 | Open weights (unrestricted) | 2025-06-10 | magistral-small-2506 | |||||
| Claude Opus 4.1 | claude-opus-4-1-20250805_16K | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Anthropic’s largest flagship model, benchmarked with up to 16,000 reasoning tokens. | Claude Opus 4.1 (16k thinking) | claude-opus-4-1-20250805-16k | |||
| Kimi K2 | kimi-k2-0711-preview | Moonshot | China | 2025-07-12 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | A preview of an updated version of Kimi K2. | kimi-k2-0711-preview | |||
| GLM 4.5 | glm-4.5 | Zhipu AI,Tsinghua University | China | 2025-08-03 | Open weights (unrestricted) | 2025-08-05 | 4.4e+24 | Zhipu AI & Tsinghua University’s large open source model. | glm-4-5 | |||
| GPT-5 Pro | gpt-5-pro-2025-10-06_high | OpenAI | United States of America | 2025-10-07 | API access | 2025-10-07 | gpt-5-pro-2025-10-06-high | |||||
| DeepSeek-R1 | DeepSeek-R1-0528 | DeepSeek-R1 (May 2025) | DeepSeek | China | 2025-05-28 | Open weights (unrestricted) | 2025-01-20 | 4.0e+24 | An updated May 2025 version of DeepSeek's reasoning model, R1. | DeepSeek-R1 (May 2025) | deepseek-r1-0528 | 139.63 |
| Claude Sonnet 4 | claude-sonnet-4-20250514_59K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 59,000 reasoning tokens. | Claude Sonnet 4 (59k thinking) | claude-sonnet-4-20250514-59k | |||
| Gemini 1.5 Flash | gemini-1.5-flash-001 | Gemini 1.5 Flash (May 2024) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-05-23 | API access | 2024-05-10 | The first version of Google DeepMind's Gemini 1.5 Flash. | gemini-1-5-flash-001 | 121.96 | ||
| Claude 2 | claude-2.0 | Anthropic | United States of America | 2023-07-11 | API access | 2023-07-11 | 3.9e+24 | Anthropic's second generation Claude model. | claude-2-0 | 119.91 | ||
| Claude 3 Opus | claude-3-opus-20240229 | Anthropic | United States of America | 2024-02-29 | API access | 2024-03-04 | The largest model in Anthropic's Claude 3 model series. | Claude 3 Opus | claude-3-opus-20240229 | 126.79 | ||
| o1-preview | o1-preview-2024-09-12 | o1-preview | OpenAI | United States of America | 2024-09-12 | API access | 2024-09-12 | The September 2024 preview version of OpenAI’s first reasoning model, o1. | o1-preview-2024-09-12 | 134.62 | ||
| Claude 3 Sonnet | claude-3-sonnet-20240229 | Anthropic | United States of America | 2024-02-29 | API access | 2024-03-04 | The mid-sized model in Anthropic's Claude 3 model series. | Claude 3 Sonnet | claude-3-sonnet-20240229 | 120.18 | ||
| GPT-4 Turbo | gpt-4-0125-preview | GPT-4 Turbo Preview (January 2024) | OpenAI | United States of America | 2024-01-25 | API access | 2023-11-06 | A January 2024 preview version of GPT-4 Turbo, OpenAI's updated GPT-4 series model. | gpt-4-0125-preview | |||
| Claude 2.1 | claude-2.1 | Anthropic | United States of America | 2023-11-21 | API access | 2023-11-21 | An updated version in Anthropic's Claude 2 series. | claude-2-1 | ||||
| Gemma 2 27B | gemma-2-27b-it | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-06-24 | Open weights (restricted use) | 2024-06-24 | 2.1e+24 | An instruction-optimized version of Google DeepMind's Gemma 2 27B. | gemma-2-27b-it | 122.91 | ||
| Claude 3 Haiku | claude-3-haiku-20240307 | Anthropic | United States of America | 2024-03-07 | API access | 2024-03-04 | The smallest model in Anthropic's Claude 3 series. | Claude 3 Haiku | claude-3-haiku-20240307 | 118.01 | ||
| GPT-4o | gpt-4o-2024-05-13 | GPT-4o (May 2024) | OpenAI | United States of America | 2024-05-13 | API access | 2024-05-13 | The first version of GPT-4o, OpenAI's last multimodal model in the GPT-4 series, which powered ChatGPT from mid-2024 to -2025. | gpt-4o-2024-05-13 | 128.42 | ||
| GPT-4 | gpt-4-0613 | GPT-4 (Jun 2023) | OpenAI | United States of America | 2023-06-13 | API access | 2023-03-15 | 2.1e+25 | The June 2023 version of GPT-4, OpenAI's 2023 multimodal flagship model. | gpt-4-0613 | 122.11 | |
| Gemma 2 9B | gemma-2-9b-it | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-06-24 | Open weights (restricted use) | 2024-06-24 | 4.3e+23 | gemma-2-9b-it | 119.44 | |||
| GPT-3.5 Turbo | gpt-3.5-turbo-1106 | GPT-3.5 Turbo (Nov 2023) | OpenAI | United States of America | 2023-11-06 | API access | 2022-11-30 | A November 2023 preview of GPT-3.5 turbo, OpenAI's updated GPT-3.5 series model. | gpt-3-5-turbo-1106 | 119.63 | ||
| Gemini 1.5 Pro | gemini-1.5-pro-001 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-05-24 | API access | 2024-02-15 | The first of two versions of a previous flagship model from Google DeepMind, Gemini 1.5 Pro. | gemini-1-5-pro-001 | 126.16 | |||
| GPT-4 Turbo | gpt-4-1106-preview | GPT-4 Turbo Preview (Nov 2023) | OpenAI | United States of America | 2023-11-06 | API access | 2023-11-06 | A November 2023 preview of GPT-4 turbo, OpenAI's updated GPT-4 series model. | gpt-4-1106-preview | |||
| GPT-4o mini | gpt-4o-mini-2024-07-18 | OpenAI | United States of America | 2024-07-18 | API access | 2024-07-18 | A smaller version of OpenAI's GPT-4o. | gpt-4o-mini-2024-07-18 | 126.10 | |||
| Gemini 1.5 Pro | gemini-1.5-pro-002 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-09-24 | API access | 2024-02-15 | The first of two versions of a previous flagship model from Google DeepMind, Gemini 1.5 Pro. | gemini-1-5-pro-002 | 132.22 | |||
| GPT-3.5 Turbo | gpt-3.5-turbo-0125 | GPT-3.5 Turbo (Jan 2024) | OpenAI | United States of America | 2024-01-25 | API access | 2022-11-30 | A January 2024 preview version of GPT-3.5 Turbo, OpenAI's updated GPT-3.5 series model. | gpt-3-5-turbo-0125 | |||
| Llama 3.1-8B | Llama-3.1-8B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 1.2e+24 | An instruction-tuned version of the 8 billion-parameter model in Meta’s LLaMA 3.1 series. | llama-3-1-8b-instruct | 118.42 | ||
| Llama 3.1-70B | Llama-3.1-70B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 7.9e+24 | An instruction-optimized version of Llama 3.1 70B, the 70 billion-parameter model in Meta's Llama 3.1 series. | llama-3-1-70b-instruct | 125.70 | ||
| Llama 3.1-405B | Llama-3.1-405B-Instruct | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 3.8e+25 | An instruction-tuned version of the 405 billion-parameter model in Meta’s LLaMA 3.1 series. | llama-3-1-405b-instruct | 128.05 | ||
| Yi-1.5-34B | Yi-1.5-34B-Chat | 01.AI | China | 2024-05-13 | Open weights (restricted use) | 2024-05-13 | 7.3e+23 | Yi-1.5-34B (chat) | yi-1-5-34b-chat | |||
| Yi-34B | Yi-34B-Chat | 01.AI | China | 2023-11-22 | Open weights (restricted use) | 2023-11-02 | 6.1e+23 | Yi-34B (chat) | yi-34b-chat | |||
| Qwen2-72B | qwen2-72b-instruct | Alibaba | China | 2024-06-07 | Open weights (unrestricted) | 2024-06-07 | 3.0e+24 | Qwen's 72 billion-parameter model in the Qwen 2 series. | Qwen2-72B | qwen2-72b-instruct | 125.47 | |
| Qwen1.5-72B | qwen1.5-72b-chat | Alibaba | China | 2024-02-04 | Open weights (restricted use) | 2024-02-04 | 1.3e+24 | Qwen1.5-72B | qwen1-5-72b-chat | |||
| Qwen1.5-32B | qwen1.5-32b-chat | Alibaba | China | 2024-04-03 | Open weights (restricted use) | 2024-02-05 | A chat-optimized 32 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-32B (chat) | qwen1-5-32b-chat | |||
| Hermes 2 Theta Llama-3 70B | Hermes-2-Theta-Llama-3-70B | Nous Research,Arcee AI | United States of America | 2024-06-20 | Open weights (restricted use) | 2024-06-20 | Nous Research’s fine-tuned Llama-3 70B optimized for instruction following and chat. | hermes-2-theta-llama-3-70b | ||||
| Llama 2-70B | Llama-2-70b-chat-hf | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | A chat-optimized version of the 70-billion parameter model in Meta's Llama 2 series. | llama-2-70b-chat-hf | |||
| Llama 3-70B | Meta-Llama-3-70B-Instruct | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.9e+24 | An instruction-tuned version of the 70 billion-parameter model in Meta’s LLaMA 3 series. | meta-llama-3-70b-instruct | 121.95 | ||
| Llama 3-8B | Meta-Llama-3-8B-Instruct | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.2e+23 | An instruction-tuned version of the 8 billion-parameter model in Meta's Llama 3 series. | meta-llama-3-8b-instruct | 116.01 | ||
| Mistral 7B | Mistral-7B-Instruct-v0.3 | Mistral AI | France | 2024-05-27 | Open weights (unrestricted) | 2023-10-10 | An instruction-tuned version of Mistral-7B-v0.3, an updated version of their 7 billion-parameter model. | mistral-7b-instruct-v0-3 | ||||
| DeepSeek LLM 67B | deepseek-llm-67b-chat | DeepSeek | China | 2023-11-29 | Open weights (restricted use) | 2024-01-05 | 8.0e+23 | DeepSeek LLM 67B (chat) | deepseek-llm-67b-chat | |||
| Mixtral 8x7B | Mixtral-8x7B-Instruct-v0.1 | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | An instruction-tuned version of Mistral's 7 billion active (56 billion total)-parameter mixture-of-experts model | mixtral-8x7b-instruct-v0-1 | |||
| Qwen2.5-72B | qwen2.5-72b-instruct | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | An instruction-tuned version of the 72 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-72B | qwen2-5-72b-instruct | 129.71 | |
| WizardLM-2 8x22B | WizardLM-2-8x22B | Microsoft | United States of America | 2024-04-15 | Open weights (unrestricted) | 2024-04-15 | WizardLM-2 8x22B | wizardlm-2-8x22b | ||||
| DBRX | dbrx-instruct | Databricks | United States of America | 2024-03-27 | Open weights (restricted use) | 2024-03-27 | 2.6e+24 | DBRX (instruct) | dbrx-instruct | |||
| Gemini 1.0 Pro | gemini-1.0-pro-001 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-02-15 | API access | 2023-12-06 | An earlier flagship model from Google DeepMind. | gemini-1-0-pro-001 | ||||
| Ministral 3B | ministral-3b-2410 | Mistral AI | France | 2024-10-16 | API access | 2024-10-16 | ministral-3b-2410 | |||||
| Mistral Large | mistral-large-2402 | Mistral AI | France | 2024-02-26 | API access | 2024-02-26 | 1.1e+25 | mistral-large-2402 | ||||
| Mixtral 8x22B | open-mixtral-8x22b | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | open-mixtral-8x22b | ||||
| Ministral 8B | ministral-8b-2410 | Mistral AI | France | 2024-10-16 | Open weights (non-commercial) | 2024-10-16 | ministral-8b-2410 | |||||
| Mistral Large 2 | mistral-large-2407 | Mistral AI | France | 2024-07-24 | Open weights (non-commercial) | 2024-07-24 | 2.1e+25 | mistral-large-2407 | 127.42 | |||
| Mixtral 8x7B | open-mixtral-8x7b | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | open-mixtral-8x7b | ||||
| Mistral 7B | open-mistral-7b | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | open-mistral-7b | |||||
| Mistral NeMo | open-mistral-nemo-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | open-mistral-nemo-2407 | |||||
| Llama 3.2 90B | Llama-3.2-90B-Vision-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | An instruction-tuned version of the 90 billion-parameter vision model in Meta's Llama 3.3 series. | llama-3-2-90b-vision-instruct | 125.56 | |||
| Tulu 3 (Tülu 3) 70B | Llama-3.1-Tulu-3-70B-DPO | Allen Institute for AI,University of Washington | United States of America | 2024-11-21 | Open weights (restricted use) | 2024-11-21 | Tülu 3 70B | llama-3-1-tulu-3-70b-dpo | ||||
| Eurus-2-7B-PRIME | Eurus-2-7B-PRIME | Tsinghua University,University of Illinois Urbana-Champaign (UIUC),Shanghai AI Lab,Peking University,Shanghai Jiao Tong University,CUHK Shenzhen Research Institute | China,United States of America | 2024-12-31 | Open weights (unrestricted) | 2025-02-03 | eurus-2-7b-prime | |||||
| Llama 3.3 70B | Llama-3.3-70B-Instruct | Meta AI | United States of America | 2024-12-06 | Open weights (restricted use) | 2024-12-06 | 6.9e+24 | An instruction-tuned version of the 70 billion-parameter model in Meta's Llama 3.3 series. | llama-3-3-70b-instruct | 127.51 | ||
| o1 | o1-2024-12-17_medium | o1 (medium) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | OpenAI's first reasoning model, evaluated with at the medium reasoning effort level. | o1-2024-12-17-medium | 140.74 | ||
| Qwen2.5-32B | qwen2.5-32b-instruct | Alibaba | China | 2024-09-17 | Open weights (unrestricted) | 2024-09-17 | 3.5e+24 | Qwen2.5-32B (instruct) | qwen2-5-32b-instruct | 128.45 | ||
| Mistral Small 3 | mistral-small-2501 | Mistral AI | France | 2025-01-25 | Open weights (unrestricted) | 2025-01-30 | 1.2e+24 | mistral-small-2501 | 126.79 | |||
| DeepSeek-R1 | DeepSeek-R1 | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-20 | 4.0e+24 | DeepSeek-R1 | deepseek-r1 | 137.97 | ||
| phi-3-medium 14B | Phi-3-medium-128k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 4.0e+23 | phi-3-medium-128k-instruct | 120.33 | |||
| Phi-4 | phi-4 | Microsoft Research | United States of America | 2024-12-12 | Open weights (unrestricted) | 2024-12-12 | 9.3e+23 | phi-4 | 130.13 | |||
| Gemini 2.0 Flash Thinking | gemini-2.0-flash-thinking-exp-01-21 | Gemini 2.0 Flash Thinking Exp | Google DeepMind,Google | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-01-21 | API access | 2024-12-19 | A January 2025 experimental version of Gemini 2.0 Flash Thinking, a small reasoning model from Google DeepMind. | gemini-2-0-flash-thinking-exp-01-21 | 135.39 | ||
| GPT-4 Turbo | gpt-4-turbo-2024-04-09 | OpenAI | United States of America | 2024-04-09 | API access | 2023-11-06 | The April 2024 version of GPT-4 Turbo, OpenAI's then-flagship language model. | gpt-4-turbo-2024-04-09 | 127.66 | |||
| GPT-4.5 | gpt-4.5-preview-2025-02-27 | GPT-4.5 Preview (Feb 2025) | OpenAI | United States of America | 2025-02-27 | API access | 2025-02-27 | 2.1e+26 | The largest model in OpenAI’s GPT series. | gpt-4-5-preview-2025-02-27 | 136.42 | |
| DeepSeek-R1-Distill-Llama-70B | DeepSeek-R1-Distill-Llama-70B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | A model based on LLaMA 70B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | deepseek-r1-distill-llama-70b | 136.37 | |||
| DeepSeek-R1-Distill-Qwen-14B | DeepSeek-R1-Distill-Qwen-14B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | A model based on Qwen 14B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | deepseek-r1-distill-qwen-14b | ||||
| Gemini 1.5 Flash 8B | gemini-1.5-flash-8b-001 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-10-03 | API access | 2024-05-10 | gemini-1-5-flash-8b-001 | |||||
| Gemma 3 27B | gemma-3-27b-it | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-03-12 | Open weights (restricted use) | 2025-03-12 | 2.3e+24 | An instruction-tuned version of the 27 billion-parameter model Google DeepMind's Gemma 3 series. | gemma-3-27b-it | 129.66 | ||
| Mistral Small 3.1 | mistral-small-2503 | Mistral AI | France | 2025-03-17 | Open weights (unrestricted) | 2025-03-17 | An updated version of Mistral small, a 24 billion-parameter model from Mistral. | mistral-small-2503 | 127.36 | |||
| Gemini 2.5 Pro | gemini-2.5-pro-exp-03-25 | Gemini 2.5 Pro Exp (Mar 2025) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-03-25 | API access | 2025-03-25 | A March 2025 preview version of Gemini 2.5 Pro. | Gemini 2.5 Pro Exp (Mar 2025) | gemini-2-5-pro-exp-03-25 | 142.98 | |
| Qwen Plus | qwen-plus-2025-01-25 | Alibaba | China | 2025-01-25 | API access | 2024-02-06 | Qwen Plus (Jan 2025) | qwen-plus-2025-01-25 | ||||
| Qwen-Turbo | qwen-turbo-2024-11-01 | Alibaba | China | 2024-11-01 | API access | 2024-02-06 | Qwen Turbo | qwen-turbo-2024-11-01 | ||||
| QWQ-Plus | qwq-plus | Alibaba | China | 2025-04-08 | API access | 2025-04-08 | QwQ-Plus | qwq-plus | ||||
| Gemini 2.5 Pro | gemini-2.5-pro-preview-05-06 | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-05-06 | API access | 2025-03-25 | A May 2025 preview version of Gemini 2.5 Pro, a flagship reasoning model from Google DeepMind. | Gemini 2.5 Pro Preview (Jun 2025) | gemini-2-5-pro-preview-05-06 | 140.76 | |
| Claude Opus 4 | claude-opus-4-20250514_16K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 16,000 reasoning tokens. | Claude Opus 4 (16k thinking) | claude-opus-4-20250514-16k | 139.85 | ||
| Claude Sonnet 4 | claude-sonnet-4-20250514_16K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 16,000 reasoning tokens. | Claude Sonnet 4 (16k thinking) | claude-sonnet-4-20250514-16k | |||
| Claude Sonnet 4 | claude-sonnet-4-20250514_32K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's mid-sized model in the Claude 4 series, benchmarked with up to 32,000 reasoning tokens. | Claude Sonnet 4 (32k thinking) | claude-sonnet-4-20250514-32k | 140.10 | ||
| Qwen3-Max | qwen3-max-2025-09-23 | Qwen3-Max-Instruct | Alibaba | China | 2025-09-24 | API access | 2025-09-05 | 1.5e+25 | A 1-trillion total parameter scale model in the Qwen 3 series. | Qwen3 Max | qwen3-max-2025-09-23 | 140.68 |
| GPT-4 | gpt-4-0314 | GPT-4 (Mar 2023) | OpenAI | United States of America | 2023-03-14 | API access | 2023-03-15 | 2.1e+25 | The March 2023 version of GPT-4, OpenAI's 2023 multimodal flagship model. | gpt-4-0314 | 125.61 | |
| gpt-oss-120b | gpt-oss-120b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | The larger one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-120b-high | |||
| gpt-oss-20b | gpt-oss-20b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 5.5e+23 | The smaller one of OpenAI's two latest open source language models, evaluated with high reasoning effort. | gpt-oss-20b-high | |||
| Gemini 2.5 Pro | gemini-2.5-pro-preview-03-25 | Gemini 2.5 Pro Preview (Mar 2025) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-04-09 | API access | 2025-03-25 | A March 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship model. | Gemini 2.5 Pro Preview (Mar 2025) | gemini-2-5-pro-preview-03-25 | 140.73 | |
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-05-20 | API access | 2025-04-17 | The May 2025 preview version of Gemini 2.5 Flash, a small reasoning model from Google DeepMind's Gemini 2.5 series. | gemini-2-5-flash-preview-05-20 | 139.27 | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-04-17 | Gemini 2.5 Flash Preview (Apr 2025) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-04-17 | API access | 2025-04-17 | An April 2025 preview of Gemini 2.5 Flash, a small model in Google DeepMind's Gemini 2.5 series. | gemini-2-5-flash-preview-04-17 | 138.48 | ||
| GLM 4.6 | glm-4.6 | Zhipu AI,Tsinghua University | China | 2025-09-30 | Open weights (unrestricted) | 2025-09-30 | 4.4e+24 | glm-4-6 | ||||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_59K | Claude Sonnet 4.5 (59k thinking) | Anthropic | United States of America | API access | 2025-09-29 | Claude Sonnet 4.5 (59k thinking) | claude-sonnet-4-5-20250929-59k | ||||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_16K | Claude Sonnet 4.5 (16k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Anthropic's latest mid-size model, part of the Claude 4.5 series, evaluated with up to 16,000 reasoning tokens. | Claude Sonnet 4.5 (16k thinking) | claude-sonnet-4-5-20250929-16k | ||
| DeepSeek-V3.1 | DeepSeek-V3.1 | DeepSeek | China | 2025-08-21 | Open weights (unrestricted) | 2025-08-21 | 3.6e+24 | An updated version of DeepSeek V3. | DeepSeek-V3.1 | deepseek-v3-1 | ||
| MPT-7B | mpt-7b | MosaicML | United States of America | 2023-05-05 | Open weights (unrestricted) | 2023-05-05 | 4.2e+22 | mpt-7b | 94.94 | |||
| chatglm2-6b | 2023-06-24 | chatglm2-6b | ||||||||||
| internlm-7b | 2023-07-05 | internlm-7b | ||||||||||
| internlm-20b | 2023-09-18 | internlm-20b | ||||||||||
| Baichuan 2-7B | Baichuan-2-7B-Base | Baichuan | China | 2023-09-20 | Open weights (restricted use) | 2023-09-20 | 1.1e+23 | baichuan-2-7b-base | 94.91 | |||
| Baichuan2-13B | Baichuan-2-13B-Base | Baichuan | China | 2023-09-06 | Open weights (restricted use) | 2023-09-06 | 2.0e+23 | baichuan-2-13b-base | 102.80 | |||
| LLaMA-7B | LLaMA-7B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-24 | 4.0e+22 | Meta's smallest model in the original Llama series. | llama-7b | 96.64 | ||
| LLaMA-13B | LLaMA-13B | Meta AI | United States of America | 2023-02-27 | Open weights (non-commercial) | 2023-02-27 | 7.8e+22 | Meta's 13 billion-parameter model in the original Llama series. | llama-13b | 101.01 | ||
| LLaMA-33B | LLaMA-33B | Meta AI | United States of America | 2023-02-27 | Open weights (non-commercial) | 2023-02-27 | 2.7e+23 | Meta's 33 billion-parameter model in the original Llama series | llama-33b | 108.72 | ||
| LLaMA-65B | LLaMA-65B | Meta AI | United States of America | 2023-02-24 | Open weights (non-commercial) | 2023-02-24 | 5.5e+23 | Meta's largest model in the original Llama series. | llama-65b | 110.89 | ||
| Llama 2-7B | Llama-2-7b | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.4e+22 | Meta's 7 billion-parameter model in the Llama 2 series of open source models. | llama-2-7b | 98.81 | ||
| Llama 2-13B | Llama-2-13b | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 1.6e+23 | Meta's 13 billion-parameter model in the Llama 2 series of open source models. | llama-2-13b | 106.48 | ||
| Llama 2-70B | Llama-2-70b-hf | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | A chat-optimized of a 70 billion-parameter model in the Llama 2 series. | llama-2-70b-hf- | 116.40 | ||
| Stable Beluga 2 | StableBeluga2 | Stability AI | United Kingdom of Great Britain and Northern Ireland | 2023-07-20 | Open weights (non-commercial) | 2023-07-20 | Stable Beluga 2 | stablebeluga2 | 117.87 | |||
| Qwen-1_8B | Qwen-1.8B | 2023-11-30 | qwen-1-8b | |||||||||
| Qwen-7B | Qwen-7B | Alibaba | China | 2023-09-28 | Open weights (restricted use) | 2023-09-28 | 1.0e+23 | Qwen's 7 billion-parameter model in the original Qwen series. | Qwen-7B | qwen-7b | 107.27 | |
| Qwen-14B | Qwen-14B | Alibaba | China | 2023-09-28 | Open weights (restricted use) | 2023-09-28 | 2.5e+23 | Qwen's 14 billion-parameter model in the original Qwen series. | Qwen-14B | qwen-14b | 112.10 | |
| InstructGPT 175B | text-davinci-001 | OpenAI | United States of America | 2022-01-27 | API access | 2022-01-27 | 3.2e+23 | text-davinci-001 | ||||
| Gopher (280B) | Gopher (280B) | DeepMind | United Kingdom of Great Britain and Northern Ireland | 2021-12-08 | Unreleased | 2021-12-08 | 6.3e+23 | A 280 billion-parameter model from DeepMind in 2021. | gopher-280b | |||
| Megatron-Turing NLG 530B | Megatron-Turing NLG 530B | Microsoft,NVIDIA | United States of America | 2022-01-28 | Unreleased | 2021-10-11 | 8.6e+23 | megatron-turing-nlg-530b | ||||
| Chinchilla | Chinchilla (70B) | DeepMind | United Kingdom of Great Britain and Northern Ireland | 2022-03-29 | Unreleased | 2022-03-29 | 5.8e+23 | Chinchilla | chinchilla-70b | |||
| PaLM (540B) | PaLM 540B | Google Research | United States of America | 2022-04-04 | Unreleased | 2022-04-04 | 2.5e+24 | Google's largest model in the PaLM series. | palm-540b | |||
| Inflection-1 | Inflection-1 | Inflection AI | United States of America | 2023-06-22 | Hosted access (no API) | 2023-06-23 | 1.0e+24 | inflection-1 | 119.25 | |||
| Falcon-7B | falcon-7b | Technology Innovation Institute | United Arab Emirates | 2023-04-24 | Open weights (unrestricted) | 2023-04-24 | 6.3e+22 | The 7 billion-parameter model in the Falcon series. | falcon-7b | 95.66 | ||
| Falcon-40B | falcon-40b | Technology Innovation Institute | United Arab Emirates | 2023-03-15 | Open weights (unrestricted) | 2023-03-15 | 2.4e+23 | The 40 billion-parameter model in the Falcon series. | falcon-40b | 104.61 | ||
| Falcon-180B | falcon-180B | Technology Innovation Institute | United Arab Emirates | 2023-09-06 | Open weights (restricted use) | 2023-09-06 | 3.8e+24 | The 180 billion-parameter model in the Falcon series. | falcon-180b | 113.60 | ||
| OpenHands + Amazon Nova Pro V1:0 | amazon.nova-pro-v1:0 | CMU | United States of America | 2024-12-03 | API access | 2024-12-03 | 6.0e+24 | amazon-nova-pro-v10 | 125.80 | |||
| Mistral 7B | Mistral-7B-Instruct-v0.2 | Mistral AI | France | Open weights (unrestricted) | 2023-10-10 | mistral-7b-instruct-v0-2 | ||||||
| Gemma 2B | gemma-2b | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-02-21 | Open weights (restricted use) | 2024-02-21 | 4.5e+22 | Google's 2 billion-parameter model in the Gemma series of open source models. | gemma-2b | 94.82 | ||
| Gemma 7B | gemma-7b | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-02-21 | Open weights (restricted use) | 2024-02-21 | 3.1e+23 | Google's 7 billion-parameter model in the Gemma series of open source models. | gemma-7b | 111.14 | ||
| Baichuan1-7B | Baichuan-7B | Baichuan | China | 2023-06-01 | Open weights (non-commercial) | 2023-06-01 | 5.0e+22 | baichuan-7b | 93.22 | |||
| phi-3.5-mini | Phi-3.5-mini-instruct | Microsoft | United States of America | 2024-08-16 | Open weights (unrestricted) | 2024-04-23 | 3.7e+22 | An instruction-tuned version of phi-3.5-mini, a small model in Microsoft's Phi 3.5 open source model series. | phi-3-5-mini-instruct | |||
| Phi-3.5-MoE | Phi-3.5-MoE-instruct | Microsoft | United States of America | 2024-08-17 | Open weights (unrestricted) | 2024-04-23 | 3.0e+23 | phi-3-5-moe-instruct | ||||
| Mistral 7B | Mistral-7B-v0.1 | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | mistral-7b-v0-1 | 111.45 | ||||
| Mistral NeMo | Mistral-Nemo-Base-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | mistral-nemo-base-2407 | |||||
| Gemma 2 9B | gemma-2-9b | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | Open weights (restricted use) | 2024-06-24 | 4.3e+23 | A 9 billion-parameter open source model from Google DeepMind in the Gemma 2 series. | gemma-2-9b | ||||
| DeepSeek-V2 (MoE-236B) | DeepSeek-V2 | DeepSeek | China | 2024-05-07 | Open weights (restricted use) | 2024-05-07 | 1.0e+24 | deepseek-v2 | 125.32 | |||
| Qwen2.5-72B | Qwen2.5-72B | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | A 72 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-72B | qwen2-5-72b | 125.76 | |
| Llama 3.1-405B | Llama-3.1-405B | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 3.8e+25 | The 405 billion-parameter model in Meta’s LLaMA 3.1 series. | llama-3-1-405b | 129.73 | ||
| DeepSeek-V3 | DeepSeek-V3-Base | DeepSeek | China | 2024-12-26 | Open weights (restricted use) | 2024-12-24 | 3.4e+24 | DeepSeek-V3 (base) | deepseek-v3-base | |||
| MPT-30B | mpt-30b | MosaicML | United States of America | 2023-06-22 | Open weights (unrestricted) | 2023-06-22 | 1.9e+23 | mpt-30b | 99.74 | |||
| Llama 2-34B | Llama-2-34b | Meta AI | United States of America | 2023-07-18 | Unreleased | 2023-07-18 | 4.1e+23 | Meta's 34 billion-parameter model in the Llama 2 series. | llama-2-34b | 106.38 | ||
| PaLM 62B | 2022-04-04 | palm-62b | ||||||||||
| Mistral 7B | Mistral-7B-Instruct-v0.1 | Mistral AI | France | 2023-09-27 | Open weights (unrestricted) | 2023-10-10 | mistral-7b-instruct-v0-1 | |||||
| Mixtral 8x7B | Mixtral-8x7B-v0.1 | Mistral AI | France | 2023-12-11 | Open weights (unrestricted) | 2023-12-11 | 7.7e+23 | mixtral-8x7b-v0-1 | 118.04 | |||
| Nemotron-4 15B | Nemotron-4 15B | NVIDIA | United States of America | 2024-02-26 | Unreleased | 2024-02-27 | 7.5e+23 | A 15 billion-parameter language model trained by Nvidia. | nemotron-4-15b | |||
| vicuna-13b-v1.1 | 2023-04-12 | vicuna-13b-v1-1 | ||||||||||
| OPT-1.3B | opt-1.3b | Meta AI | United States of America | 2022-05-11 | Open weights (non-commercial) | 2022-06-21 | opt-1-3b | |||||
| GPT-Neo-2.7B | gpt-neo-2.7B | EleutherAI | United States of America | 2023-03-30 | Open weights (unrestricted) | 2021-03-21 | 7.9e+21 | gpt-neo-2-7b | ||||
| GPT-2 (1.5B) | gpt2-xl | OpenAI | United States of America | 2019-11-05 | Open weights (unrestricted) | 2019-02-14 | 1.9e+21 | The largest version of GPT-2, OpenAI's second generation transformer-based large pretrained language model. | gpt2-xl | |||
| XGen-7B | xgen-7b-8k-base | Salesforce | United States of America | 2023-06-27 | Open weights (unrestricted) | 2023-09-07 | 8.0e+22 | XGen-7B | xgen-7b-8k-base | |||
| open_llama_7b | 2023-06-07 | open-llama-7b | ||||||||||
| RedPajama-INCITE-7B-Base | 2023-05-04 | redpajama-incite-7b-base | ||||||||||
| GPT-NeoX-20B | gpt-neox-20b | EleutherAI | United States of America | 2022-04-07 | Open weights (unrestricted) | 2022-02-09 | 9.3e+22 | gpt-neox-20b | ||||
| opt-13b | 2022-05-11 | opt-13b | ||||||||||
| GPT-J-6B | gpt-j-6b | EleutherAI,LAION | Germany,United States of America | 2021-08-05 | Open weights (unrestricted) | 2021-05-01 | 1.5e+22 | gpt-j-6b | ||||
| Dolly 2.0-12b | dolly-v2-12b | Databricks | United States of America | 2023-04-11 | Open weights (unrestricted) | 2023-04-12 | dolly-v2-12b | |||||
| Cerebras-GPT-13B | Cerebras-GPT-13B | Cerebras Systems | United States of America | 2023-03-20 | Open weights (unrestricted) | 2023-04-06 | 2.3e+22 | cerebras-gpt-13b | ||||
| stablelm-tuned-alpha-7b | 2023-04-19 | stablelm-tuned-alpha-7b | ||||||||||
| Yi 6B | Yi-6B | 01.AI | China | 2023-11-02 | Open weights (restricted use) | 2023-11-02 | 1.3e+23 | Yi-6B (chat) | yi-6b | 103.63 | ||
| Yi-34B | Yi-34B | 01.AI | China | 2023-11-02 | Open weights (restricted use) | 2023-11-02 | 6.1e+23 | Yi-34B | yi-34b | |||
| GPT-3.5 Turbo | gpt-3.5-turbo-0613 | GPT-3.5 Turbo (June 2023) | OpenAI | United States of America | 2023-06-13 | API access | 2022-11-30 | GPT-3.5 Turbo (June 2023) | gpt-3-5-turbo-0613 | |||
| Baichuan 1-13B | Baichuan-13B-Base | Baichuan | China | 2023-07-11 | Open weights (restricted use) | 2023-07-11 | 9.4e+22 | baichuan-13b-base | ||||
| phi-3-mini 3.8B | Phi-3-mini-4k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 7.5e+22 | A 4 billion-parameter model in Microsoft's Phi-3 open source model series with a 4000-token context length. | phi-3-mini-4k-instruct | 118.81 | ||
| phi-3-small 7.4B | Phi-3-small-8k-instruct | Microsoft | United States of America | 2024-04-23 | Open weights (unrestricted) | 2024-04-23 | 2.1e+23 | A 7 billion-parameter model in Microsoft's Phi-3 open source model series with a 8000-token context length. | phi-3-small-8k-instruct | 121.10 | ||
| Phi-2 | phi-2 | Microsoft | United States of America | 2023-12-12 | Open weights (unrestricted) | 2023-12-12 | 2.3e+22 | phi-2 | 111.98 | |||
| Llama 2-13B | Llama-2-13b-chat | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 1.6e+23 | A chat-optimized version of Meta's 13 billion-parameter model in the Llama 2 series. | llama-2-13b-chat | |||
| Llama 2-70B | Llama-2-70b-chat | Meta AI | United States of America | 2023-07-18 | Open weights (restricted use) | 2023-07-18 | 8.1e+23 | A chat-optimized version of Meta's 70 billion-parameter model in the Llama 2 series. | llama-2-70b-chat | |||
| Baichuan2-13B-Chat | 2023-09-06 | baichuan2-13b-chat | ||||||||||
| Qwen-14B | Qwen-14B-Chat | Alibaba | China | 2023-09-24 | Open weights (restricted use) | 2023-09-28 | 2.5e+23 | A chat-optimized version of Qwen's first-generation 14 billion-parameter model. | Qwen-14B (chat) | qwen-14b-chat | ||
| internlm-chat-20b | 2023-09-17 | internlm-chat-20b | ||||||||||
| Yi 6B | Yi-6B-Chat | 01.AI | China | 2023-11-22 | Open weights (restricted use) | 2023-11-02 | 1.3e+23 | Yi-6B | yi-6b-chat | |||
| Yi-9B | 2024-03-01 | yi-9b | ||||||||||
| PaLM 2-S | 2023-05-17 | palm-2-s | ||||||||||
| PaLM 2-M | 2023-05-17 | palm-2-m | ||||||||||
| PaLM 2-L | 2023-05-17 | palm-2-l | ||||||||||
| Qwen2.5-Coder-0.5B | 2024-09-18 | qwen2-5-coder-0-5b | ||||||||||
| DeepSeek Coder 1.3B | deepseek-coder-1.3b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 1.6e+22 | DeepSeek Coder 1.3B (base) | deepseek-coder-1-3b-base | |||
| Qwen2.5-Coder (1.5B) | Qwen2.5-Coder-1.5B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-09-18 | 5.1e+22 | Qwen's 1.5 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-1.5B | qwen2-5-coder-1-5b | 104.50 | |
| StarCoder 2 3B | starcoder2-3b | Hugging Face,ServiceNow,NVIDIA,BigCode | 2024-02-22 | Open weights (restricted use) | 2024-02-29 | 5.9e+22 | StarCoder 2 3B | starcoder2-3b | ||||
| Qwen2.5-Coder-3B | 2024-09-18 | qwen2-5-coder-3b | ||||||||||
| StarCoder 2 7B | starcoder2-7b | Hugging Face,ServiceNow,NVIDIA,BigCode | 2024-02-20 | Open weights (restricted use) | 2024-02-29 | 1.5e+23 | StarCoder 2 7B | starcoder2-7b | ||||
| DeepSeek Coder 6.7B | deepseek-coder-6.7b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 8.0e+22 | DeepSeek Coder 6.7B (base) | deepseek-coder-6-7b-base | |||
| DeepSeek-Coder-V2-Lite-Base | 2024-06-13 | deepseek-coder-v2-lite-base | ||||||||||
| CodeQwen1.5-7B | 2024-04-15 | codeqwen1-5-7b | ||||||||||
| Qwen2.5-Coder (7B) | Qwen2.5-Coder-7B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-09-18 | 2.5e+23 | Qwen's 7 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-7B | qwen2-5-coder-7b | 113.01 | |
| StarCoder 2 15B | starcoder2-15b | Hugging Face,ServiceNow,NVIDIA,BigCode | 2024-02-20 | Open weights (restricted use) | 2024-02-29 | 3.9e+23 | StarCoder 2 15B | starcoder2-15b | ||||
| Qwen2.5-Coder-14B | 2024-09-18 | qwen2-5-coder-14b | ||||||||||
| DeepSeek Coder 33B | deepseek-coder-33b-base | DeepSeek,Peking University | China | 2023-11-02 | Open weights (restricted use) | 2024-01-25 | 4.0e+23 | DeepSeek Coder 33B (base) | deepseek-coder-33b-base | |||
| DeepSeek-Coder-V2 236B | DeepSeek-Coder-V2-Base | DeepSeek | China | 2024-06-17 | Open weights (restricted use) | 2024-06-17 | 1.3e+24 | DeepSeek-Coder-V2 (base) | deepseek-coder-v2-base | |||
| Qwen2.5-Coder (32B) | Qwen2.5-Coder-32B | Alibaba | China | 2024-09-18 | Open weights (unrestricted) | 2024-11-12 | 1.1e+24 | Qwen's 32 billion-parameter model in the Qwen 2.5 Coder series of open source models. | Qwen2.5-Coder-32B | qwen2-5-coder-32b | 117.98 | |
| Phi-1.5 | phi-1_5 | Microsoft | United States of America | 2023-09-11 | Open weights (unrestricted) | 2023-09-11 | 1.2e+21 | phi-1-5 | 93.32 | |||
| GPT-3.5 | text-davinci-002 | OpenAI | United States of America | 2022-03-15 | API access | 2022-11-28 | 2.6e+24 | text-davinci-002 | ||||
| Claude Instant | claude-instant-1.1 | Anthropic | United States of America | API access | 2023-08-09 | claude-instant-1-1 | ||||||
| Claude Instant | claude-instant-1.2 | Anthropic | United States of America | 2023-08-09 | API access | 2023-08-09 | claude-instant-1-2 | |||||
| T5-Large | t5-large | |||||||||||
| T5-11B | T5-11B | United States of America | Open weights (unrestricted) | 2019-10-23 | 3.3e+22 | The largest, 11 billion-parameter version of T5, an early large language model from Google. | t5-11b | |||||
| T5-3B | T5-3B | United States of America | Open weights (unrestricted) | 2019-10-23 | 9.0e+21 | The 3 billion-parameter version of T5, an early large language model from Google. | t5-3b | |||||
| Mixtral 8x22B | Mixtral-8x22B-Instruct-v0.1 | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | An instruction-tuned version of Mistral's 8 billion active (176 billion total)-parameter mixture-of-experts model | mixtral-8x22b-instruct-v0-1 | |||
| Gemini 2.0 Flash | gemini-2.0-flash-02-05 | Google DeepMind,Google | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-02-05 | API access | 2024-12-11 | A February 2025 version of Gemini 2.0 Flash, a small language model from Google DeepMind. | gemini-2-0-flash-02-05 | ||||
| GLaM | GLaM (MoE) | United States of America | 2021-12-13 | Unreleased | 2021-12-13 | 3.6e+23 | glam-moe | |||||
| Amazon Nova Lite | amazon.nova-lite-v1:0 | Amazon | United States of America | 2024-12-03 | API access | 2024-12-03 | amazon-nova-lite-v10 | |||||
| Amazon Nova Micro | amazon.nova-micro-v1:0 | Amazon | United States of America | 2024-12-03 | API access | 2024-12-03 | amazon-nova-micro-v10 | |||||
| Claude 1.3 | claude-1.3 | Anthropic | United States of America | 2023-04-18 | API access | 2023-04-18 | An updated version of Anthropic's first generation Claude model. | claude-1-3 | ||||
| c4ai-command-r-08-2024 | 2024-08-30 | c4ai-command-r-08-2024 | ||||||||||
| Command R+ | c4ai-command-r-plus-08-2024 | Cohere,Cohere for AI | Canada | 2024-08-30 | Open weights (non-commercial) | 2024-04-04 | Command R+ | c4ai-command-r-plus-08-2024 | ||||
| Falcon 2 11B | falcon-11b | Falcon 2-11B | Technology Innovation Institute | United Arab Emirates | 2024-05-09 | Open weights (restricted use) | 2024-05-09 | 3.6e+23 | falcon-11b | |||
| Gemini 1.5 Flash | gemini-1.5-flash-0514 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-05-14 | API access | 2024-05-10 | gemini-1-5-flash-0514 | |||||
| Gemini 2.0 Flash | gemini-2.0-flash-exp | Google DeepMind,Google | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-12-11 | API access | 2024-12-11 | A December 2024 experimental version of Gemini 2.0 Flash, a small langauge model from Google DeepMind. | gemini-2-0-flash-exp | 130.91 | |||
| GPT-4 Turbo | gpt-4-turbo | OpenAI | United States of America | 2023-11-06 | API access | 2023-11-06 | An updated version of GPT-4 and OpenAI's then-new flagship language model. | gpt-4-turbo | ||||
| Llama 3.2 11B | Llama-3.2-11B-Vision-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 5.8e+23 | An 11 billion-parameter instruction-tuned vision language model in Meta's Llama 3.2 series. | llama-3-2-11b-vision-instruct | |||
| mistral-small-2402 | 2024-02-26 | mistral-small-2402 | ||||||||||
| Mixtral 8x22B | Mixtral-8x22B-v0.1 | Mistral AI | France | 2024-04-17 | Open weights (unrestricted) | 2024-04-17 | 2.3e+24 | mixtral-8x22b-v0-1 | ||||
| Qwen1.5-7B | Qwen1.5-7B | Alibaba | China | 2024-02-04 | Open weights (unrestricted) | 2024-02-04 | 1.7e+23 | A 7 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-7B | qwen1-5-7b | ||
| Qwen1.5-14B | qwen1.5-14B | Alibaba | China | 2024-02-04 | Open weights (unrestricted) | 2024-02-04 | 3.4e+23 | A 14 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-14B | qwen1-5-14b | ||
| Qwen1.5-32B | qwen1.5-32B | Alibaba | China | 2024-02-04 | Open weights (restricted use) | 2024-02-05 | A 32 billion-parameter model in the Qwen 1.5 series. | Qwen1.5-32B | qwen1-5-32b | |||
| Qwen2.5-14B | qwen2.5-14b-instruct | Alibaba | China | 2025-02-26 | Open weights (unrestricted) | 2024-09-19 | 1.6e+24 | An instruction-tuned version of Qwen 2.5 14b. | Qwen2.5-14B (instruct) | qwen2-5-14b-instruct | ||
| Qwen2.5-7B | qwen2.5-7b-instruct | Alibaba | China | 2025-02-26 | Open weights (unrestricted) | 2024-09-19 | 8.2e+23 | An instruction-tuned version of Qwen 2.5 7b. | Qwen2.5-7B | qwen2-5-7b-instruct | ||
| Yi-Large | Yi-large | 01.AI | China | 2024-05-13 | API access | 2024-05-13 | 1.8e+24 | Yi-Large | yi-large | |||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_2K | Claude Sonnet 4.5 (2k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Claude Sonnet 4.5 (2k thinking) | claude-sonnet-4-5-20250929-2k | |||
| GPT-5 | gpt-5-2025-08-07_low | GPT-5 (low) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | OpenAI's latest flagship frontier model, benchmarked using low reasoning effort. | GPT-5 (low) | gpt-5-2025-08-07-low | ||
| Claude Opus 4 | claude-opus-4-20250514_2K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 2,000 reasoning tokens. | Claude Opus 4 (2k thinking) | claude-opus-4-20250514-2k | |||
| GPT-5 | gpt-5-2025-08-07_minimal | GPT-5 (minimal) | OpenAI | United States of America | 2025-08-07 | API access | 2025-08-07 | OpenAI's latest flagship frontier model, benchmarked using minimal reasoning effort. | GPT-5 (minimal) | gpt-5-2025-08-07-minimal | ||
| Claude Sonnet 4 | claude-sonnet-4-20250514_2K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's mid-sized Claude 4 series model, evaluated at up to 2,000 reasoning tokens. | Claude Sonnet 4 (2k thinking) | claude-sonnet-4-20250514-2k | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_2K | Claude 3.7 Sonnet (2k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | A previous flagship reasoning model from Anthropic, benchmarked with up to 2,000 reasoning tokens allowed. | Claude 3.7 Sonnet (2k thinking) | claude-3-7-sonnet-20250219-2k | |
| o3-pro | o3-pro-2025-06-10_high | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with high reasoning effort. | o3-pro-2025-06-10-high | |||||
| Claude Opus 4 | claude-opus-4-20250514_8K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 8,000 reasoning tokens. | Claude Opus 4 (8k thinking) | claude-opus-4-20250514-8k | |||
| Gemini 2.5 Pro | gemini-2.5-pro-preview-06-05_1K | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-06-05 | API access | 2025-03-25 | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model, evaluated with up to 1,000 reasoning tokens. | Gemini 2.5 Pro Preview (Jun 2025, 1k thinking) | gemini-2-5-pro-preview-06-05-1k | ||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-04-17 (24K thinking) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-04-17 | API access | 2025-04-17 | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 24,000 reasoning tokens. | gemini-2-5-flash-preview-04-17-24k-thinking | ||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_1K | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-05-20 | API access | 2025-04-17 | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 1,000 reasoning tokens. | gemini-2-5-flash-preview-05-20-1k | ||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_8K | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-05-20 | API access | 2025-04-17 | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 8,000 reasoning tokens. | gemini-2-5-flash-preview-05-20-8k | ||||
| Claude Sonnet 4 | claude-sonnet-4-20250514_8K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Claude Sonnet 4 (8k thinking) | claude-sonnet-4-20250514-8k | ||||
| o3-pro | o3-pro-2025-06-10_low | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with low reasoning effort. | o3-pro-2025-06-10-low | |||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_16K | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-05-20 | API access | 2025-04-17 | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 16,000 reasoning tokens. | gemini-2-5-flash-preview-05-20-16k | ||||
| o3-pro | o3-pro-2025-06-10_medium | OpenAI | United States of America | 2025-06-10 | 2025-06-10 | An enhanced version of OpenAI's previous flagship reasoning model o3, evaluated with medium reasoning effort. | o3-pro-2025-06-10-medium | |||||
| codex-mini-2025-05-16 | 2025-05-16 | codex-mini-2025-05-16 | ||||||||||
| o1 | o1-pro-2025-03-19_low | OpenAI | United States of America | 2025-03-19 | API access | 2024-12-05 | An An enhanced version of OpenAI's first reasoning model,o1, evaluated with low reasoning effort. | o1-pro-2025-03-19-low | ||||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_8K | Claude 3.7 Sonnet (8k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | A previous flagship reasoning model from Anthropic, benchmarked with up to 9,000 reasoning tokens allowed. | Claude 3.7 Sonnet (8k thinking) | claude-3-7-sonnet-20250219-8k | |
| Claude Sonnet 4 | claude-sonnet-4-20250514_1K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's mid-sized Claude 4 series model, evaluated at up to 1,000 reasoning tokens. | Claude Sonnet 4 (1k thinking) | claude-sonnet-4-20250514-1k | |||
| o1 | o1-2024-12-17_low | o1 (low) | OpenAI | United States of America | 2024-12-17 | API access | 2024-12-05 | OpenAI's first reasoning model, evaluated with at the low reasoning tokens level. | o1-2024-12-17-low | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_1K | Claude 3.7 Sonnet (1k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | A previous flagship reasoning model from Anthropic, benchmarked with up to 1,000 reasoning tokens allowed. | Claude 3.7 Sonnet (1k thinking) | claude-3-7-sonnet-20250219-1k | |
| Llama 4 Maverick | Llama-4-Maverick-17B-128E-Instruct | Llama 4 Maverick | Meta AI | United States of America | 2025-04-05 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | The 17 billion active (400 billion total)-parameter model in Meta's Llama 4 series. | llama-4-maverick-17b-128e-instruct | 129.31 | |
| o3-mini | o3-mini-2025-01-31_low | OpenAI | United States of America | 2025-01-31 | API access | 2025-01-31 | A smaller version of OpenAI's o3 reasoning model, evaluated with low reasoning effort. | o3-mini-2025-01-31-low | ||||
| Claude Opus 4 | claude-opus-4-20250514_1K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 1,000 reasoning tokens. | Claude Opus 4 (1k thinking) | claude-opus-4-20250514-1k | |||
| Grok 3 | grok-3 | xAI | United States of America | 2025-04-09 | API access | 2025-02-17 | 3.5e+26 | XAI's third generation flagship model. | Grok 3 | grok-3 | ||
| Gemini 2.5 Pro | gemini-2.5-pro-preview-06-05_32K | Gemini 2.5 Pro Preview (Jun 2025) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | API access | 2025-03-25 | A June 2025 preview version of Gemini 2.5 Pro, Google DeepMind's flagship frontier model, evaluated with up to 32,000 reasoning tokens. | Gemini 2.5 Pro Preview (Jun 2025, 32k thinking) | gemini-2-5-pro-preview-06-05-32k | |||
| Claude Opus 4 | claude-opus-4-20250514_32K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 32,000 reasoning tokens. | Claude Opus 4 (32k thinking) | claude-opus-4-20250514-32k | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-05-20_23K | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-05-20 | API access | 2025-04-17 | A May 2025 preview version of Gemini 2.5 Flash, a small model in Google DeepMind's Gemini 2.5 series, evaluated with up to 23,000 reasoning tokens. | gemini-2-5-flash-preview-05-20-23k | ||||
| GPT-4o | chatgpt-4o-03-27-2025 | OpenAI | United States of America | 2025-03-27 | API access | 2024-05-13 | A version of GPT-4o that was behind the ChatGPT interface, released in March 2025. | ChatGPT-4o (Mar 2025) | chatgpt-4o-03-27-2025 | |||
| Qwen3-32B | Qwen3-32B | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 7.1e+24 | The 32 billion-parameter model in the Qwen 3 series. | Qwen3 32B | qwen3-32b | ||
| Gemini 2.0 Flash | gemini-exp-1206 | Google DeepMind,Google | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-12-06 | API access | 2024-12-11 | A December 2024 experimental version of Gemini 2.0 Flash, a small langauge model from Google DeepMind. | gemini-exp-1206 | ||||
| GPT-4o | chatgpt-4o-01-29-2025 | OpenAI | United States of America | 2025-01-29 | API access | 2024-05-13 | A version of GPT-4o that was behind the ChatGPT interface, released in January 2025. | ChatGPT-4o (Jan 2025) | chatgpt-4o-01-29-2025 | |||
| QwQ-32B | QwQ-32B | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | QwQ 32B | qwq-32b | |||
| DeepSeek-V2.5 | DeepSeek-V2.5 | DeepSeek | China | 2024-09-06 | Open weights (restricted use) | 2024-09-06 | 1.8e+24 | DeepSeek-V2.5 (Sept 2024) | deepseek-v2-5 | |||
| Qwen2.5-Coder (32B) | Qwen2.5-Coder-32B-Instruct | Alibaba | China | 2024-11-21 | Open weights (unrestricted) | 2024-11-12 | 1.1e+24 | An instruction-tuned and coding-optimized 32 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-Coder-32B (instruct) | qwen2-5-coder-32b-instruct | ||
| Yi-Lightning | yi-lightning | 01.AI | China | 2024-12-02 | API access | 2024-10-18 | 1.5e+24 | Yi-Lightning | yi-lightning | |||
| Cohere Command A | c4ai-command-a-03-2025 | Cohere | Canada | 2025-03-13 | Open weights (non-commercial) | 2025-03-13 | Command A | c4ai-command-a-03-2025 | ||||
| Codestral | codestral-2501 | Mistral AI | France | 2025-01-13 | Open weights (non-commercial) | 2024-05-29 | A January 2025 version of a 2024 coding-optimized model from Mistral. | codestral-2501 | ||||
| openhands-lm-32b-v0.1 | 2024-03-26 | openhands-lm-32b-v0-1 | ||||||||||
| GPT-3.5 | text-davinci-003 | OpenAI | United States of America | 2022-11-28 | API access | 2022-11-28 | 2.6e+24 | text-davinci-003 | ||||
| GPT-4 | gpt-4-32k-0314 | OpenAI | United States of America | 2023-03-14 | API access | 2023-03-15 | 2.1e+25 | gpt-4-32k-0314 | ||||
| OPT-175B | opt-175b | Meta AI | United States of America | 2022-05-02 | Open weights (non-commercial) | 2022-05-02 | 4.3e+23 | Facebook's 66 billion-parameter model in the OPT series. | opt-175b | |||
| GPT-3 175B (davinci) | davinci | OpenAI | United States of America | API access | 2020-05-28 | 3.1e+23 | GPT-3 175B (davinci) | davinci | ||||
| OPT-66B | opt-66b | Meta AI | United States of America | 2022-05-03 | Open weights (non-commercial) | 2022-06-21 | 1.1e+23 | Facebook's 66 billion-parameter model in the OPT series. | opt-66b | |||
| BLOOM-176B | bloom | Hugging Face,BigScience | France,United States of America | 2022-07-06 | Open weights (restricted use) | 2022-07-11 | 3.7e+23 | bloom | ||||
| GPT-3 6.7B | curie | OpenAI | United States of America | 2020-06-22 | 1.2e+22 | curie | ||||||
| InstructGPT 6B | text-curie-001 | OpenAI | United States of America | API access | 2022-01-27 | text-curie-001 | ||||||
| InstructGPT 1.3B | text-babbage-001 | OpenAI | United States of America | API access | 2022-01-27 | text-babbage-001 | ||||||
| GPT-3 XL | babbage | OpenAI | United States of America | 2020-06-22 | 2.4e+21 | babbage | ||||||
| GPT-3 Medium | ada | OpenAI | United States of America | 2020-06-22 | 6.4e+20 | ada | ||||||
| InstructGPT 350M | text-ada-001 | OpenAI | United States of America | 2022-01-27 | text-ada-001 | |||||||
| ml-elephant | Zoo Ml-Ephant | ml-elephant | ||||||||||
| QwQ-32B | QwQ-32B-Preview-Q8_0-GGUF | Alibaba | China | 2024-12-27 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | QwQ 32B Preview (Q8 GGUF) | qwq-32b-preview-q8-0-gguf | |||
| Llama 3.3 70B | llama3.3:70b-instruct-q8_0 | Meta AI | United States of America | 2024-12-06 | Open weights (restricted use) | 2024-12-06 | 6.9e+24 | llama3-370b-instruct-q8-0 | ||||
| Llama 3.1-405B | llama3.1:405b-instruct-q4_K_M | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 3.8e+25 | llama3-1405b-instruct-q4-k-m | ||||
| Phi-4 | phi4:14b-q8_0 | Microsoft Research | United States of America | 2025-02-26 | Open weights (unrestricted) | 2024-12-12 | 9.3e+23 | phi414b-q8-0 | ||||
| Llama 3.1-8B | llama3.1:8b-instruct-q8_0 | Meta AI | United States of America | 2024-07-23 | Open weights (restricted use) | 2024-07-23 | 1.2e+24 | llama3-18b-instruct-q8-0 | ||||
| Gemma 2 27B | gemma2:27b-instruct-q8_0 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-06-27 | Open weights (restricted use) | 2024-06-24 | 2.1e+24 | gemma227b-instruct-q8-0 | ||||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-04-17 (16K thinking) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-04-17 | API access | 2025-04-17 | An April 2025 preview version of Gemini 2.5 Flash, a small reasoning model in the Gemini 2.5 series from DeepMind, evaluated with up to 16,000 reasoning tokens. | gemini-2-5-flash-preview-04-17-16k-thinking | ||||
| Qwen3-30B-A3B | Qwen3-30B-A3B | Alibaba | China | 2025-04-29 | Open weights (unrestricted) | 2025-04-29 | 6.5e+23 | Qwen3-30B-A3B | qwen3-30b-a3b | |||
| Llama 3-8B | Meta-Llama-3-8B | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.2e+23 | Meta's 8 billion-parameter model in the Llama 3 series. | meta-llama-3-8b | |||
| Llama 3-70B | Meta-Llama-3-70B | Meta AI | United States of America | 2024-04-18 | Open weights (restricted use) | 2024-04-18 | 7.9e+24 | Meta's 70 billion-parameter model in the Llama 3 series. | meta-llama-3-70b | |||
| Reka Flash 3 | reka-flash-3 | Reka AI | United States of America | 2025-03-10 | Open weights (unrestricted) | 2025-03-10 | Reka Flash 3 | reka-flash-3 | ||||
| DeepSeek-R1-Distill-Qwen-32B | DeepSeek-R1-Distill-Qwen-32B | DeepSeek | China | 2025-01-20 | Open weights (unrestricted) | 2025-01-22 | A model based on Qwen 32B and additionally trained on outputs from DeepSeek’s R1 model, published as part of DeepSeek R1’s release. | deepseek-r1-distill-qwen-32b | ||||
| Mistral NeMo | Mistral-Nemo-Instruct-2407 | Mistral AI | France | 2024-07-18 | Open weights (unrestricted) | 2024-07-18 | mistral-nemo-instruct-2407 | |||||
| Qwen2-VL-72B-Instruct | 2024-08-29 | qwen2-vl-72b-instruct | ||||||||||
| Llama 3.2 3B | Llama-3.2-3B-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 1.7e+23 | A 3 billion-parameter instruction-tuned model in Meta's Llama 3.2 series. | llama-3-2-3b-instruct | |||
| Llama 3.2 1B | Llama-3.2-1B-Instruct | Meta AI | United States of America | 2024-09-24 | Open weights (restricted use) | 2024-09-24 | 6.6e+22 | A 1 billion-parameter instruction-tuned model in Meta's Llama 3.2 series. | llama-3-2-1b-instruct | |||
| Qwen2-VL-7B-Instruct | 2024-08-29 | qwen2-vl-7b-instruct | ||||||||||
| Gemini 2.5 Flash | gemini-2.5-flash | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-06-17 | API access | 2025-04-17 | A small model from Google DeepMind's Gemini 2.5 series. | Gemini 2.5 Flash (Jun 2025) | gemini-2-5-flash | |||
| Claude Opus 4 | claude-opus-4-20250514_12K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's largest model in the Claude 4 series (superseded by Claude Opus 4.1), benchmarked with up to 12,000 reasoning tokens. | Claude Opus 4 (12k thinking) | claude-opus-4-20250514-12k | |||
| Claude Sonnet 4.5 | claude-sonnet-4-5-20250929_12K | Claude Sonnet 4.5 (12k thinking) | Anthropic | United States of America | 2025-09-29 | API access | 2025-09-29 | Claude Sonnet 4.5 (12k thinking) | claude-sonnet-4-5-20250929-12k | |||
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_12K | Claude 3.7 Sonnet (12k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | A previous flagship reasoning model from Anthropic, benchmarked with up to 12,000 reasoning tokens allowed. | Claude 3.7 Sonnet (12k thinking) | claude-3-7-sonnet-20250219-12k | |
| Claude Sonnet 4 | claude-sonnet-4-20250514_12K | Anthropic | United States of America | 2025-05-22 | API access | 2025-05-22 | Anthropic's mid-sized Claude 4 series model, evaluated at up to 12,000 reasoning tokens. | Claude Sonnet 4 (12k thinking) | claude-sonnet-4-20250514-12k | |||
| gpt-oss-120b | gpt-oss-120b | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | The larger one of OpenAI's two latest open source language models. | gpt-oss-120b | |||
| GLM-4.5-Air | 2025-07-20 | glm-4-5-air | ||||||||||
| Qwen3-Coder-480B-A35B | Qwen3-Coder-480B-A35B-Instruct | Alibaba | China | 2025-07-31 | Open weights (unrestricted) | 2025-07-22 | 1.6e+24 | Qwen3 Coder | qwen3-coder-480b-a35b-instruct | |||
| DeepSeek-V3.1 | DeepSeek-V3.1_thinking | DeepSeek | China | 2025-08-21 | Open weights (unrestricted) | 2025-08-21 | 3.6e+24 | DeepSeek-V3.1 (thinking) | deepseek-v3-1-thinking | |||
| Qwen3-235B-A22B | Qwen3-235B-A22B-Instruct-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | A 22 billion active (235 billion total)-parameter mixture-of-experts instruction-tuned model in the Qwen 3 series. | Qwen3 Non-thinking (Jul 2025) | qwen3-235b-a22b-instruct-2507 | ||
| qwen3-coder-plus | 2025-09-23 | qwen3-coder-plus | ||||||||||
| grok-code-fast-1 | 2025-08-28 | grok-code-fast-1 | ||||||||||
| MiniMax-M1-80k | MiniMax-M1-80k | MiniMax | China | 2025-06-13 | Open weights (unrestricted) | 2025-06-13 | 4.3e+24 | minimax-m1-80k | ||||
| Qwen2.5-Max | qwen2.5-max | Alibaba | China | 2025-01-28 | API access | 2025-01-28 | The largest model in the Qwen 2.5 series. | Qwen2.5-Max | qwen2-5-max | |||
| computer-use-preview-2025-03-11 | 2025-03-11 | computer-use-preview-2025-03-11 | ||||||||||
| Qwen2.5-72B | Qwen2.5-VL-72B-Instruct | Alibaba | China | 2024-09-19 | Open weights (unrestricted) | 2024-09-19 | 7.8e+24 | A 72 billion-parameter instruction-tuned vision language model in the Qwen 2.5 series. | Qwen2.5-VL-72B | qwen2-5-vl-72b-instruct | 129.98 | |
| Claude 3.7 Sonnet | claude-3-7-sonnet-20250219_15K | Claude 3.7 Sonnet (15k thinking) | Anthropic | United States of America | 2025-02-24 | API access | 2025-02-24 | 3.3e+25 | A previous flagship reasoning model from Anthropic, benchmarked with up to 15,000 reasoning tokens allowed. | Claude 3.7 Sonnet (15k thinking) | claude-3-7-sonnet-20250219-15k | |
| Pixtral 12B | Pixtral-12B-2409 | Mistral AI | France | 2024-09-17 | Open weights (unrestricted) | 2024-09-17 | pixtral-12b-2409 | |||||
| GPT-5 | gpt-5-codex | OpenAI | United States of America | 2025-09-15 | API access | 2025-08-07 | GPT-5-codex | gpt-5-codex | ||||
| glaive-swe-v1 | glaive-swe-v1 | |||||||||||
| Grok 4 Fast | grok-4-fast | xAI | United States of America | 2025-09-19 | API access | 2025-09-19 | Grok 4 Fast | grok-4-fast | ||||
| MPT-30B | mpt-30b-instruct | MosaicML | United States of America | 2023-06-22 | Open weights (unrestricted) | 2023-06-22 | 1.9e+23 | mpt-30b-instruct | ||||
| Vicuna-13B-v1.3 | vicuna-13b-v1.3 | Large Model Systems Organization,University of California (UC) Berkeley | United States of America | 2023-06-18 | Open weights (restricted use) | 2023-06-22 | Vicuna-13B-v1.3 | vicuna-13b-v1-3 | ||||
| T5-Small | t5-small | |||||||||||
| T5-Base | t5-base | |||||||||||
| Switch-Base | switch-base | |||||||||||
| Switch-Large | switch-large | |||||||||||
| Claude Opus 4.1 | claude-opus-4-1-20250805_32K | Anthropic | United States of America | 2025-08-05 | API access | 2025-08-05 | Anthropic’s largest flagship model, benchmarked with up to 32,000 reasoning tokens. | Claude Opus 4.1 (32k thinking) | claude-opus-4-1-20250805-32k | |||
| GPT-3 175B (davinci) | davinci-002 | OpenAI | United States of America | API access | 2020-05-28 | 3.1e+23 | GPT-3 175B (davinci-002) | davinci-002 | ||||
| GPT-3.5 Turbo | gpt-3.5-turbo-instruct | OpenAI | United States of America | 2023-09-18 | API access | 2022-11-30 | An instruction-tuned version of OpenAI's GPT-3.5 turbo. | gpt-3-5-turbo-instruct | ||||
| GPT-3.5 | code-davinci-002 | OpenAI | United States of America | API access | 2022-11-28 | 2.6e+24 | code-davinci-002 | |||||
| Falcon-40B | falcon-40b-instruct | Technology Innovation Institute | United Arab Emirates | 2023-05-25 | Open weights (unrestricted) | 2023-03-15 | 2.4e+23 | falcon-40b-instruct | ||||
| DeepSeek-Coder-V2-Lite-Instruct | 2024-06-13 | deepseek-coder-v2-lite-instruct | ||||||||||
| DeepSeek-Coder-V2 236B | DeepSeek-Coder-V2-Instruct | DeepSeek | China | 2024-06-17 | Open weights (restricted use) | 2024-06-17 | 1.3e+24 | An instruction-tuned version of DeepSeek Coder V2, a 236 billion-parameter coding-optimized model from DeepSeek. | DeepSeek-Coder-V2 (instruct) | deepseek-coder-v2-instruct | ||
| Qwen2.5-Coder-3B-Instruct | 2024-11-06 | qwen2-5-coder-3b-instruct | ||||||||||
| Qwen2.5-Coder (7B) | Qwen2.5-Coder-7B-Instruct | Alibaba | China | 2024-09-17 | Open weights (unrestricted) | 2024-09-18 | 2.5e+23 | An instruction-tuned and coding-optimized 7 billion-parameter model in the Qwen 2.5 series. | Qwen2.5-Coder-7B (instruct) | qwen2-5-coder-7b-instruct | ||
| Qwen2.5-Coder-14B-Instruct | 2024-11-06 | qwen2-5-coder-14b-instruct | ||||||||||
| QwQ-32B | QwQ-32B (16K thinking) | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | QwQ 32B (16k thinking) | qwq-32b-16k-thinking | |||
| Gemini 2.5 Flash | gemini-2.5-flash-preview-09-25 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-09-25 | API access | 2025-04-17 | Gemini 2.5 Flash (Sep 2025) | gemini-2-5-flash-preview-09-25 | ||||
| gemini-robotics-er-1.5-preview | 2025-09-26 | gemini-robotics-er-1-5-preview | ||||||||||
| Gemini 2.5 Flash-Lite | gemini-2.5-flash-lite-preview-09-2025 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-09-25 | API access | 2025-06-15 | Gemini 2.5 Flash-Lite (Sep 2025) | gemini-2-5-flash-lite-preview-09-2025 | ||||
| Computer-Using Agent (CUA) | CUA | OpenAI | United States of America | 2025-01-23 | Hosted access (no API) | 2025-01-23 | cua | |||||
| GPT-4V | gpt-4-1106-vision-preview | OpenAI | United States of America | 2023-11-06 | API access | 2023-09-25 | gpt-4-1106-vision-preview | |||||
| Gemini 2.0 Flash-Lite | gemini-2.0-flash-lite | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-02-05 | API access | 2024-02-05 | The smallest model in Google DeepMind's Gemini 2.0 series. | gemini-2-0-flash-lite | ||||
| Gemini 2.0 Flash-Lite | gemini-2.0-flash-lite-preview-02-05 | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-02-05 | API access | 2024-02-05 | A February 2025 preview version of the smallest model in Google DeepMind's Gemini 2.0 series. | gemini-2-0-flash-lite-preview-02-05 | ||||
| Dracarys2-72B-Instruct | 2024-09-30 | dracarys2-72b-instruct | ||||||||||
| learnlm-1.5-pro-experimental | learnlm-1-5-pro-experimental | |||||||||||
| sonar | Perplexity Sonar | sonar | ||||||||||
| Dracarys2-Llama-3.1-70B-Instruct | 2024-08-14 | dracarys2-llama-3-1-70b-instruct | ||||||||||
| QwQ-32B | QwQ-32B-Preview | Alibaba | China | 2024-11-28 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | QwQ 32B Preview | qwq-32b-preview | |||
| OLMo 2 Furious 13B | OLMo-2-1124-13B-Instruct | Allen Institute for AI,University of Washington,New York University (NYU) | United States of America | 2024-12-31 | Open weights (unrestricted) | 2024-12-31 | 4.6e+23 | olmo-2-1124-13b-instruct | ||||
| BLIP-2 (Q-Former) | blip2-opt-2.7b | Salesforce Research | United States of America | 2023-02-06 | Open weights (unrestricted) | 2023-01-30 | 1.2e+21 | blip2-opt-2-7b | ||||
| Phi-3.5-vision-instruct | 2024-08-16 | phi-3-5-vision-instruct | ||||||||||
| MM1-3B-Chat | 2024-03-14 | mm1-3b-chat | ||||||||||
| MM1-7B-Chat | 2024-03-14 | mm1-7b-chat | ||||||||||
| llava-v1.6-vicuna-7b | 2024-01-31 | llava-v1-6-vicuna-7b | ||||||||||
| llama3-llava-next-8b | 2024-04-20 | llama3-llava-next-8b | ||||||||||
| Qwen-VL-Chat | 2023-08-20 | qwen-vl-chat | ||||||||||
| Gemini 1.0 Pro Vision | gemini-1.0-pro-vision | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2024-01-04 | API access | 2024-02-15 | A vision language model from Google DeepMind. | gemini-1-0-pro-vision | ||||
| falcon-11B-vlm | 2024-05-21 | falcon-11b-vlm | ||||||||||
| llava-v1.6-vicuna-13b | 2024-01-31 | llava-v1-6-vicuna-13b | ||||||||||
| llava-v1.6-mistral-7b | 2024-01-31 | llava-v1-6-mistral-7b | ||||||||||
| instructblip-vicuna-7b | 2023-05-22 | instructblip-vicuna-7b | ||||||||||
| instructblip-vicuna-13b | 2023-12-25 | instructblip-vicuna-13b | ||||||||||
| InternVL-Chat-ViT-6B-Vicuna-7B | 2023-12-25 | internvl-chat-vit-6b-vicuna-7b | ||||||||||
| InternVL-Chat-ViT-6B-Vicuna-13B | 2024-08-16 | internvl-chat-vit-6b-vicuna-13b | ||||||||||
| llava-v1.5-7b | 2023-10-05 | llava-v1-5-7b | ||||||||||
| QwQ-32B | chutes/QwQ-32B | Alibaba | China | 2025-03-05 | Open weights (unrestricted) | 2025-03-06 | 3.5e+24 | QwQ 32B (Chutes) | chutes-qwq-32b | |||
| Qwen3-235B-A22B | chutes/Qwen3-235B-A22B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Qwen3 (Apr 2025) (Chutes) | chutes-qwen3-235b-a22b | |||
| Qwen3-32B | chutes/Qwen3-32B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 7.1e+24 | Qwen3 32B (Chutes) | chutes-qwen3-32b | |||
| Qwen3-30B-A3B | chutes/Qwen3-30B-A3B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 6.5e+23 | Qwen3-30B-A3B (Chutes) | chutes-qwen3-30b-a3b | |||
| Qwen3-14B | chutes/Qwen3-14B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 3.2e+24 | Qwen3 14B | chutes-qwen3-14b | |||
| Qwen3-8B | chutes/Qwen3-8B | Alibaba | China | 2025-04-28 | Open weights (unrestricted) | 2025-04-29 | 1.8e+24 | Qwen3 8B (Chutes) | chutes-qwen3-8b | |||
| Grok-3 mini | grok-3-mini-beta_medium | xAI | United States of America | 2025-04-09 | API access | 2025-02-19 | A beta release of a smaller version of XAI's third generation Grok model, evaluated with medium reasoning effort. | grok-3-mini-beta-medium | ||||
| DeepSeek-V3 | chutes/DeepSeek-V3-0324 | DeepSeek-V3 (Mar 2025) | DeepSeek | China | 2025-03-24 | Open weights (restricted use) | 2024-12-24 | 3.4e+24 | DeepSeek-V3 (Mar 2025) (Chutes) | chutes-deepseek-v3-0324 | ||
| Gemma 3 27B | chutes/Gemma-3-27b-It | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-03-11 | Open weights (restricted use) | 2025-03-12 | 2.3e+24 | Gemma-3-27b-it (Chutes) | chutes-gemma-3-27b-it | |||
| Llama 4 Maverick | chutes/Llama-4-Maverick-17B-128E-Instruct | Meta AI | United States of America | 2025-04-06 | Open weights (restricted use) | 2025-04-05 | 2.2e+24 | Llama 4 Maverick (128 experts) (Chutes) | chutes-llama-4-maverick-17b-128e-instruct | |||
| Llama 4 Scout | chutes/Llama-4-Scout-17B-16E Instruct | Meta AI | United States of America | 2025-04-06 | Open weights (restricted use) | 2025-04-05 | 4.1e+24 | Llama 4 Maverick (16 experts) (Chutes) | chutes-llama-4-scout-17b-16e-instruct | |||
| DeepSeek-R1 | chutes/DeepSeek-R1-0528 | DeepSeek-R1 (May 2025) | DeepSeek | China | 2025-05-28 | Open weights (unrestricted) | 2025-01-20 | 4.0e+24 | DeepSeek-R1 (May 2025) (Chutes) | chutes-deepseek-r1-0528 | ||
| Gemini 2.5 Flash-Lite | gemini-2.5-flash-lite-preview-06-17-thinking | gemini-2.5-flash-lite-preview-06-17 (Default thinking length) | Google DeepMind | United Kingdom of Great Britain and Northern Ireland,United States of America | 2025-06-17 | API access | 2025-06-15 | A June 2025 preview of Gemini 2.5 Flash Lite. | Gemini 2.5 Flash-Lite (Jun 2025) | gemini-2-5-flash-lite-preview-06-17-thinking | ||
| chutes/GLM-4.5-FP8 | 2025-07-27 | chutes-glm-4-5-fp8 | ||||||||||
| Qwen3-235B-A22B | chutes/Qwen3-235B-A22B-Thinking-2507 | Alibaba | China | 2025-07-25 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Qwen3 (Jul 2025) (Chutes) | chutes-qwen3-235b-a22b-thinking-2507 | |||
| gpt-oss-120b | chutes/gpt-oss-120b_high | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | gpt-oss-120b (high) (Chutes) | chutes-gpt-oss-120b-high | |||
| chutes/Qwen3-Next-80B-A3B-Instruct | 2025-09-11 | chutes-qwen3-next-80b-a3b-instruct | ||||||||||
| gpt-oss-120b | chutes/gpt-oss-120b | OpenAI | United States of America | 2025-08-05 | Open weights (unrestricted) | 2025-08-05 | 4.9e+24 | gpt-oss-120b (Chutes) | chutes-gpt-oss-120b | |||
| DeepSeek-V3.2-Exp | 2025-09-29 | deepseek-v3-2-exp | ||||||||||
| Kimi K2 | fireworks/Kimi-K2-Instruct-0905 | Moonshot | China | 2025-09-05 | Open weights (restricted use) | 2025-07-11 | 3.0e+24 | fireworks-kimi-k2-instruct-0905 | ||||
| Qwen3-235B-A22B | parasail-qwen3-235b-a22b-instruct-2507 | Alibaba | China | 2025-09-01 | Open weights (unrestricted) | 2025-04-29 | 4.8e+24 | Qwen3 (Parasail) | parasail-qwen3-235b-a22b-instruct-2507 | |||
| nvidia-nemotron-nano-9b-v2 | 2025-08-18 | nvidia-nemotron-nano-9b-v2 | ||||||||||
| deepinfra/Qwen3-Next-80B-A3B-Instruct | deepinfra-qwen3-next-80b-a3b-instruct |
However, the gap varies considerably over time, sometimes even closing completely. Until the release of o1-mini, Llama 3.1-405B was rated on par with the closed-source state-of-the-art model, Claude 3.5 Sonnet.
You can see more detailed analysis about the gap in our earlier article.
Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.
Learn more about this graph
We calculate the average gap between closed-weight and open-weight state-of-the-art performance according to our internal capability metric, the Epoch Capabilities Index (ECI). ECI is a composite measure which captures performance across many benchmarks.
Analysis
To calculate the average time gap, we calculate the horizontal distance between the two lines across the range of ECI values where such lines can be drawn. Since ECI is estimated with some noise, we count an open-source model as “catching up” to a previous SOTA if their difference in scores is not statistically significant. For example, DeepSeek-V2 is counted as having caught up to GPT-4 after about 14 months, despite getting an ECI score of 125 vs. GPT-4’s 126.
We use a similar procedure to estimate the average ECI gap, taking the vertical distance across all dates where such a vertical line can be drawn.
We find an average “horizontal” time gap of 3.5 months, with a 90% confidence interval of 1.1 to 5.3 months. Along the vertical dimension, we find an average gap of 7 ECI points, with a 90% confidence interval of 0 to 14 units.
We also note that the current gap likely appears larger than it really is; we do not yet have enough evaluations of frontier open-weight models like gpt-oss-120b or MiniMax-M2 to assign ECI scores, but it is likely that they improve on DeepSeek R1 (May 2025).
Code for our analysis is available here.
