Spending on training large-scale ML models is growing at a rate of 2.4x per year. The most advanced models now cost hundreds of millions of dollars, with expenses measured by amortizing cluster costs over the training period. About half of this spending is on GPUs, with the remainder on other hardware and energy.
| Model | Domain | Training compute (FLOP) | Training compute cost (2023 USD) | Organization | Publication date |
|---|---|---|---|---|---|
| GNMT | Language | 6.6e+21 | $200k | 2016-09-26 | |
| Xception | Vision | 4.4e+20 | $10k | 2016-10-07 | |
| PolyNet | Vision | 6.4e+19 | $600 | Chinese University of Hong Kong (CUHK) | 2016-11-17 |
| MoE-Multi | Language | 9.4e+19 | $4k | Jagiellonian University, Google Brain | 2017-01-23 |
| JFT | Vision | 8.4e+20 | $20k | Google Research, Carnegie Mellon University (CMU) | 2017-07-10 |
| AlphaGo Zero | Games | 6.5e+20 | $600k | DeepMind | 2017-10-18 |
| AlphaGo Master | Games | 3.4e+20 | $500k | DeepMind | 2017-10-19 |
| ResNeXt-101 32x48d | Vision | 8.7e+21 | $100k | 2018-05-02 | |
| Big Transformer for Back-Translation | Language | 4.8e+20 | $2k | Facebook AI Research, Google Brain | 2018-08-28 |
| GPT-2 (1.5B) | Language | 1.9e+21 | $4k | OpenAI | 2019-02-14 |
| XLNet | Language | 6.2e+21 | $10k | Carnegie Mellon University (CMU), Google Brain | 2019-06-01 |
| RoBERTa Large | Language | 8.5e+21 | $90k | Facebook, University of Washington | 2019-07-01 |
| Megatron-BERT | Language | 2.2e+22 | $200k | NVIDIA | 2019-09-17 |
| Megatron-LM (8.3B) | Language | 9.1e+21 | $100k | NVIDIA | 2019-09-17 |
| T5-11B | Language | 3.3e+22 | $80k | 2019-10-23 | |
| T5-3B | Language | 9.0e+21 | $20k | 2019-10-23 | |
| AlphaStar | Games | 1.1e+23 | $100k | DeepMind | 2019-10-30 |
| XLM-RoBERTa | Language | 2.1e+22 | $80k | Facebook AI | 2019-11-05 |
| Noisy Student (L2) | Vision | 2.6e+22 | $40k | Carnegie Mellon University (CMU), Google | 2019-11-11 |
| OpenAI Five | Games | 6.7e+22 | $4M | OpenAI | 2019-12-13 |
| OpenAI Five Rerun | Games | 1.3e+22 | $300k | OpenAI | 2019-12-13 |
| Meena | Language | 1.1e+23 | $200k | Google Brain | 2020-01-28 |
| Turing-NLG | Language | 1.6e+22 | $50k | Microsoft | 2020-02-13 |
| GPT-3 175B (davinci) | Language | 3.1e+23 | $2M | OpenAI | 2020-05-28 |
| GShard (dense) | Language | 4.8e+22 | $300k | 2020-06-30 | |
| LUKE | Language | 1.8e+22 | $4k | University of Washington, National Institute of Informatics | 2020-10-02 |
| DALL-E | Image generation | 4.7e+22 | $100k | OpenAI | 2021-01-05 |
| Switch | Language | 8.2e+22 | $100k | 2021-01-11 | |
| Meta Pseudo Labels | Vision | 4.8e+22 | $50k | Google Brain, Google AI | 2021-03-01 |
| PLUG | Language | 3.6e+22 | $100k | Alibaba | 2021-04-19 |
| ProtBERT-BFD | Biology | 3.9e+22 | $50k | Technical University of Munich, NVIDIA, Seoul National University, Google, Oak Ridge National Laboratory, Med AI Technology | 2021-05-04 |
| ByT5-XXL | Language | 8.1e+22 | $90k | Google, Google Research | 2021-05-28 |
| ViT-G/14 | Vision | 5.8e+22 | $4k | Google Brain, Google Research | 2021-06-08 |
| Jurassic-1-Jumbo | Language | 3.7e+23 | $800k | AI21 Labs | 2021-08-11 |
| FLAN 137B | Language | 2.0e+24 | $200k | Google Research | 2021-09-03 |
| HyperCLOVA 204B | Language | 2.0e+23 | $400k | NAVER | 2021-09-10 |
| Megatron-Turing NLG 530B | Language | 8.6e+23 | $4M | Microsoft, NVIDIA | 2021-10-11 |
| Gopher (280B) | Language | 6.3e+23 | $600k | DeepMind | 2021-12-08 |
| GLaM | Language | 3.6e+23 | $500k | 2021-12-13 | |
| LaMDA | Language | 3.6e+23 | $200k | 2022-02-10 | |
| PaLM (540B) | Language | 2.5e+24 | $3M | Google Research | 2022-04-04 |
| OPT-175B | Language | 4.3e+23 | $700k | Meta AI | 2022-05-02 |
| Parti | Image generation | 5.1e+23 | $400k | Google Research | 2022-06-22 |
| GPT-3.5 | Language | 2.6e+24 | $5M | OpenAI | 2022-11-28 |
| GPT-4 | Multimodal, Language, Vision | 2.1e+25 | $40M | OpenAI | 2023-03-15 |
| PaLM 2 | Language | 7.3e+24 | $5M | 2023-05-10 | |
| InternLM | Language | 1.0e+24 | $2M | Shanghai AI Lab, SenseTime | 2023-07-06 |
| Falcon-180B | Language | 3.8e+24 | $10M | Technology Innovation Institute | 2023-09-06 |
| Amazon Titan | Language, Image generation | 4.8e+24 | $8M | Amazon | 2023-09-28 |
| Inflection-2 | Language | 1.0e+25 | $10M | Inflection AI | 2023-11-22 |
| Gemini 1.0 Ultra | Multimodal, Language, Vision | 5.0e+25 | $30M | Google DeepMind | 2023-12-06 |
| Gemini 1.5 Pro | Language, Multimodal | 1.6e+25 | $8M | Google DeepMind | 2024-02-15 |
| MegaScale (Production) | Language | 3.9e+24 | $3M | ByteDance, Peking University | 2024-02-23 |
| Mistral Large | Language | 1.1e+25 | $10M | Mistral AI | 2024-02-26 |
| Inflection-2.5 | Language | 8.0e+24 | $10M | Inflection AI | 2024-03-07 |
| Nemotron-4 340B | Language | 1.8e+25 | $20M | NVIDIA | 2024-06-14 |
| Llama 3.1-405B | Language | 3.8e+25 | $50M | Meta AI | 2024-07-23 |
| Grok-2 | Language, Vision, Multimodal | 3.0e+25 | $30M | xAI | 2024-08-13 |
| Grok 4 | Language, Multimodal, Vision, Speech | 5.0e+26 | $500M | xAI | 2025-07-09 |
Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.
