Since 2010, the training compute used to create AI models has been growing at a rate of 4.4x per year. Most of this growth comes from increased spending, although improvements in hardware have also played a role.
| Model | Domain | Training compute (FLOP) | Organization | Publication date |
|---|---|---|---|---|
| Feedforward NN | Vision | 3.5e+14 | University of Montreal / Université de Montréal | 2010-05-13 |
| iCCCP | Vision | 1.1e+15 | Massachusetts Institute of Technology (MIT) | 2010-06-13 |
| Pooling CNN (NORB) | Vision | 1.5e+15 | University of Bonn | 2010-09-15 |
| Pooling CNN (Caltech 101) | Vision | 1.2e+15 | University of Bonn | 2010-09-15 |
| RNN LM | Language | 5.4e+16 | Johns Hopkins University | 2010-09-26 |
| Deep Autoencoders | Vision | 3.7e+16 | University of Toronto | 2011-04-29 |
| High Performance CNN (NORB) | Vision | 2.6e+16 | IDSIA, SUPSI | 2011-07-16 |
| CNN Committee (MNIST) | Vision | 5.2e+16 | IDSIA | 2011-09-18 |
| CNN Committee (NIST) | Vision | 2.6e+16 | IDSIA | 2011-09-18 |
| CNN committee (traffic sign) | Vision | 9.9e+14 | IDSIA | 2011-10-03 |
| Dropout (CIFAR) | Vision | 4.3e+15 | University of Toronto | 2012-06-03 |
| Dropout (ImageNet) | Vision | 2.7e+17 | University of Toronto | 2012-06-03 |
| Dropout (MNIST) | Vision | 6.0e+15 | University of Toronto | 2012-06-03 |
| Unsupervised High-level Feature Learner | Vision | 6.0e+17 | 2012-07-12 | |
| LSTM LM | Language | 1.7e+16 | RWTH Aachen University | 2012-09-09 |
| AlexNet | Vision | 4.7e+17 | University of Toronto | 2012-09-30 |
| DistBelief Speech | Speech | 3.1e+17 | 2012-12-03 | |
| DNN EM segmentation | Vision | 4.8e+17 | IDSIA, SUPSI | 2012-12-03 |
| DistBelief NNLM | Language | 2.6e+18 | 2013-01-16 | |
| ReLU-Speech | Speech | 1.3e+17 | Google, University of Toronto, New York University (NYU) | 2013-05-26 |
Hierarchical Scene Labeling (Stanford Background) | Vision | 2.4e+17 | New York University (NYU) | 2013-08-01 |
| RCTM | Language | 9.3e+15 | University of Oxford | 2013-10-01 |
| RNTN | Language | 1.4e+16 | Stanford University | 2013-10-01 |
| Word2Vec (large) | Language | 3.9e+16 | 2013-10-16 | |
| Visualizing CNNs | Vision | 5.3e+17 | New York University (NYU) | 2013-11-12 |
| TransE | Language | 1.3e+18 | Universite de Technologie de Compiègne – CNRS, Google | 2013-12-05 |
| DQN | Games | 2.8e+15 | DeepMind | 2013-12-19 |
| Image generation | Vision | 4.8e+14 | University of Amsterdam | 2013-12-20 |
| GANs | Image generation | 5.2e+17 | University of Montreal / Université de Montréal | 2014-06-10 |
| SPPNet | Vision | 3.4e+18 | Microsoft, Xi’an Jiaotong University, University of Science and Technology of China (USTC) | 2014-06-18 |
| SmooCT | Games | 6.9e+16 | University College London (UCL) | 2014-07-01 |
| ACF-WIDER | Vision | 7.6e+13 | Chinese Academy of Sciences | 2014-07-15 |
| RNNsearch-50* | Language | 1.6e+18 | Jacobs University Bremen, University of Montreal / Université de Montréal | 2014-09-01 |
| VGG19 | Vision | 1.1e+19 | University of Oxford | 2014-09-04 |
| VGG16 | Vision | 1.2e+19 | University of Oxford | 2014-09-04 |
| Seq2Seq LSTM | Language | 5.6e+19 | 2014-09-10 | |
| SPN-4+KN5 | Language | 4.4e+16 | Singapore University of Technology & Design, DSO National Laboratories | 2014-09-14 |
| GoogLeNet / InceptionV1 | Vision | 1.5e+18 | Google, University of Michigan, University of North Carolina | 2014-09-17 |
| TA-CNN | Vision | 1.1e+16 | Chinese University of Hong Kong (CUHK) | 2014-11-29 |
| SNM-skip | Language | 3.0e+20 | 2014-12-03 | |
| Fractional Max-Pooling | Vision | 1.0e+17 | University of Warwick | 2014-12-18 |
| ADAM (CIFAR-10) | Vision | 6.2e+14 | University of Amsterdam, OpenAI, University of Toronto | 2014-12-22 |
| MSRA (C, PReLU) | Vision | 2.4e+19 | Microsoft Research | 2015-02-06 |
| genCNN + dyn eval | Language | 3.4e+16 | Chinese Academy of Sciences, Huawei Noah's Ark Lab, Dublin City University | 2015-03-17 |
| TC-DNN-BLSTM-DNN | Speech | 1.9e+17 | Carnegie Mellon University (CMU) | 2015-04-06 |
| U-Net | Vision | 5.1e+16 | University of Freiburg | 2015-05-18 |
| DCNN | Vision | 4.8e+17 | University of Maryland, Rutgers University | 2015-08-07 |
| AlphaGo Fan | Games | 3.8e+20 | DeepMind | 2015-10-01 |
| SAF R-CNN | Vision | 1.2e+19 | Beijing Institute of Technology, Sun Yat-sen University, Panasonic R&D, National University of Singapore | 2015-10-28 |
| Inception v3 | Vision | 1.0e+20 | Google, University College London (UCL) | 2015-12-02 |
| ResNet-101 (ImageNet) | Vision | 7.0e+18 | Microsoft | 2015-12-10 |
| ResNet-152 (ImageNet) | Vision | 1.0e+19 | Microsoft | 2015-12-10 |
| Variational (untied weights, MC) LSTM (Large) | Language | 5.9e+15 | University of Cambridge | 2015-12-16 |
| AlphaGo Lee | Games | 1.9e+21 | DeepMind | 2016-01-27 |
| Named Entity Recognition model | Language | 9.7e+16 | Carnegie Mellon University (CMU) | 2016-03-04 |
| R-FCN | Vision | 7.2e+17 | Tsinghua University, Microsoft Research | 2016-06-21 |
| ResNet-200 | Vision | 3.0e+19 | Microsoft Research Asia | 2016-09-17 |
| GNMT | Language | 6.6e+21 | 2016-09-26 | |
| Pointer Sentinel-LSTM (medium) | Language | 7.5e+15 | MetaMind Inc, Salesforce | 2016-09-26 |
| Xception | Vision | 4.4e+20 | 2016-10-07 | |
| SPIDER2 | Biology | 1.8e+16 | Griffith University, University of Iowa, Dezhou University | 2016-10-28 |
| VD-LSTM+REAL Large | Language | 2.1e+16 | Salesforce Research, Stanford University | 2016-11-04 |
| NASv3 (CIFAR-10) | Vision | 2.2e+21 | Google Brain | 2016-11-05 |
| NAS with base 8 and shared embeddings | Language | 1.0e+16 | Google Brain | 2016-11-05 |
| BIDAF | Language | 3.5e+18 | University of Washington, Allen Institute for AI | 2016-11-05 |
| PolyNet | Vision | 6.4e+19 | Chinese University of Hong Kong (CUHK) | 2016-11-17 |
| HR-ResNet101 | Vision | 7.1e+18 | Carnegie Mellon University (CMU) | 2016-12-13 |
| EnhanceNet | Vision | 1.3e+17 | Max Planck Institute for Intelligent Systems | 2016-12-23 |
| DeepStack | Games | 1.4e+19 | University of Alberta, Charles University, Czech Technical University | 2017-01-06 |
| MoE-Multi | Language | 9.4e+19 | Jagiellonian University, Google Brain | 2017-01-23 |
| Transformer | Language | 7.4e+18 | Google Research, Google Brain | 2017-06-12 |
| DeepLoc | Biology | 5.8e+17 | Technical University of Denmark, University of Copenhagen | 2017-07-07 |
| JFT | Vision | 8.4e+20 | Google Research, Carnegie Mellon University (CMU) | 2017-07-10 |
| ConvS2S (ensemble of 8 models) | Language | 5.6e+19 | Meta AI | 2017-07-25 |
AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (WT2) | Language | 3.0e+17 | Salesforce Research | 2017-08-07 |
| RetinaNet-R101 | Vision | 2.1e+18 | Facebook AI Research | 2017-08-07 |
| OpenAI TI7 DOTA 1v1 | Games | 6.0e+20 | OpenAI | 2017-08-11 |
| EI-REHN-1000D | Language | 1.1e+16 | Korea Advanced Institute of Science and Technology (KAIST) | 2017-08-14 |
| Libratus | Games | 5.5e+20 | Carnegie Mellon University (CMU) | 2017-08-19 |
GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (WT2) | Language | 4.6e+17 | Ben-Gurion University of the Negev | 2017-08-29 |
| PyramidNet | Vision | 2.3e+15 | Korea Advanced Institute of Science and Technology (KAIST) | 2017-09-06 |
| ISS | Language | 3.4e+15 | Duke University, Microsoft | 2017-09-15 |
| AWD-LSTM+WT+Cache+IOG (WT2) | Language | 3.2e+15 | NTT Communication Science Laboratories | 2017-09-26 |
| AlphaGo Zero | Games | 6.5e+20 | DeepMind | 2017-10-18 |
| AlphaGo Master | Games | 3.4e+20 | DeepMind | 2017-10-19 |
| Fraternal dropout + AWD-LSTM 3-layer (WT2) | Language | 3.1e+17 | Jagiellonian University, Mila - Quebec AI (originally Montreal Institute for Learning Algorithms), University of Montreal / Université de Montréal | 2017-10-31 |
| AWD-LSTM-MoS + dynamic evaluation (WT2, 2017) | Language | 3.4e+18 | Carnegie Mellon University (CMU) | 2017-11-10 |
| AlphaZero | Games | 1.1e+20 | DeepMind | 2017-12-05 |
| ELMo | Language | 3.3e+15 | University of Washington, Allen Institute for AI | 2018-02-01 |
| QRNN | Language | 6.9e+17 | Salesforce Research | 2018-02-01 |
| IMPALA | Games | 1.7e+20 | DeepMind | 2018-02-05 |
| 4 layer QRNN (h=2500) | Language | 5.9e+17 | Salesforce Research | 2018-03-22 |
| YOLOv3 | Vision | 1.3e+19 | University of Washington | 2018-04-08 |
| ResNeXt-101 32x48d | Vision | 8.7e+21 | 2018-05-02 | |
| Dropout-LSTM+Noise(Bernoulli) (WT2) | Language | 1.3e+17 | Columbia University, New York University (NYU), Princeton University | 2018-05-03 |
| aLSTM(depth-2)+RecurrentPolicy (WT2) | Language | 7.3e+16 | University of Manchester, Alan Turing Institute | 2018-05-22 |
| GPT-1 | Language | 1.8e+19 | OpenAI | 2018-06-01 |
| FTW (For The Win) | Games | 3.5e+19 | DeepMind | 2018-07-03 |
| Big-Little Net | Vision | 2.5e+17 | IBM | 2018-07-10 |
| Big-Little Net (speech) | Speech | 4.3e+17 | IBM | 2018-07-10 |
| Big Transformer for Back-Translation | Language | 4.8e+20 | Facebook AI Research, Google Brain | 2018-08-28 |
| (ensemble): AWD-LSTM-DOC (fin) × 5 (WT2) | Language | 6.7e+17 | NTT Communication Science Laboratories, Tohoku University | 2018-08-30 |
| Transformer + Simple Recurrent Unit | Language | 1.1e+19 | ASAPP, Cornell University, Google, Princeton University | 2018-09-17 |
| LSTM+NeuralCache | Language | 9.8e+14 | KU Leuven, ESAT - PSI, Apple | 2018-09-24 |
| Transformer (Adaptive Input Embeddings) WT103 | Language | 4.5e+19 | Facebook AI Research | 2018-09-28 |
| BERT-Large | Language | 2.8e+20 | 2018-10-11 | |
| TrellisNet | Language | 2.8e+18 | Carnegie Mellon University (CMU), Bosch Center for Artificial Intelligence, Intel Labs | 2018-10-15 |
Mesh-TensorFlow Transformer 2.9B (translation) | Language | 6.8e+19 | Google Brain | 2018-11-05 |
| Mesh-TensorFlow Transformer 4.9B (language) | Language | 1.6e+20 | Google Brain | 2018-11-05 |
| Fine-tuned-AWD-LSTM-DOC (fin) | Language | 5.2e+16 | Samsung R&D Institute Russia | 2018-11-12 |
| Multi-cell LSTM | Language | 2.0e+15 | University of Hyderabad | 2018-11-15 |
| StyleGAN | Image generation | 3.9e+16 | NVIDIA | 2018-12-12 |
| Transformer-XL (257M) | Language | 3.8e+20 | Carnegie Mellon University (CMU), Google Brain | 2019-01-09 |
| Hanabi 4 player | Games | 4.3e+18 | DeepMind, University of Oxford, Carnegie Mellon University (CMU), Google Brain | 2019-02-01 |
| GPT-2 (1.5B) | Language | 1.9e+21 | OpenAI | 2019-02-14 |
| KataGo | Games | 2.3e+19 | Jane Street | 2019-02-27 |
| SciBERT | Language | 8.9e+19 | Allen Institute for AI | 2019-03-26 |
| Cross-lingual alignment | Language | 2.6e+18 | Tel Aviv University, Massachusetts Institute of Technology (MIT) | 2019-04-04 |
| WeNet (Penn Treebank) | Language | 7.3e+17 | Amazon | 2019-04-08 |
| BERT-Large-CAS (PTB+WT2+WT103) | Language | 1.5e+20 | Amazon | 2019-04-20 |
| MuseNet | Audio | 2.2e+20 | OpenAI | 2019-04-25 |
| AWD-LSTM-DRILL + dynamic evaluation† (WT2) | Language | 4.1e+17 | IDIAP | 2019-05-14 |
| DLRM-2020 | Recommendation | 4.0e+18 | Facebook AI | 2019-05-31 |
| XLNet | Language | 6.2e+21 | Carnegie Mellon University (CMU), Google Brain | 2019-06-01 |
| Transformer-XL Large + Phrase Induction | Language | 3.8e+20 | Massachusetts Institute of Technology (MIT), University of Illinois Urbana-Champaign (UIUC) | 2019-06-04 |
| AWD-LSTM + MoS + Partial Shuffled | Language | 3.2e+17 | University of Texas at Austin | 2019-06-10 |
| RoBERTa Large | Language | 8.5e+21 | Facebook, University of Washington | 2019-07-01 |
| Pluribus | Games | 6.6e+16 | Facebook AI Research | 2019-07-11 |
| trRosetta | Biology | 3.8e+19 | Nankai University, University of Washington, Tianjin University, Harvard University | 2019-08-22 |
| UDSMProt | Biology | 6.4e+17 | Fraunhofer Heinrich Hertz Institute | 2019-09-04 |
| Megatron-LM (1.2B) | Language | 1.1e+22 | NVIDIA | 2019-09-17 |
| Megatron-LM (8.3B) | Language | 9.1e+21 | NVIDIA | 2019-09-17 |
| Megatron-BERT | Language | 2.2e+22 | NVIDIA | 2019-09-17 |
| AlphaX-1 | Vision | 8.9e+17 | Facebook AI Research, Brown University | 2019-10-02 |
| DistilBERT | Language | 1.2e+19 | Hugging Face | 2019-10-02 |
| T5-3B | Language | 9.0e+21 | 2019-10-23 | |
| T5-11B | Language | 3.3e+22 | 2019-10-23 | |
| AlphaStar | Games | 1.1e+23 | DeepMind | 2019-10-30 |
| Base LM + kNN LM + Continuous Cache | Language | 3.1e+19 | Stanford University, Facebook AI Research | 2019-11-01 |
| XLM-RoBERTa | Language | 2.1e+22 | Facebook AI | 2019-11-05 |
| CamemBERT | Language | 8.3e+20 | Facebook, INRIA, Sorbonne University | 2019-11-10 |
| Sandwich Transformer | Language | 2.4e+19 | Allen Institute for AI, Facebook AI Research | 2019-11-10 |
| Noisy Student (L2) | Vision | 2.6e+22 | Carnegie Mellon University (CMU), Google | 2019-11-11 |
| MuZero | Games | 4.8e+19 | DeepMind | 2019-11-19 |
| Transformer-XL DeFINE (141M) | Language | 1.7e+18 | University of Washington, Allen Institute for AI | 2019-11-27 |
| MMLSTM (WT-2) | Language | 1.9e+17 | Beijing University of Posts and Telecommunications, University of West London | 2019-12-05 |
| MMLSTM (PTB) | Language | 5.8e+16 | Beijing University of Posts and Telecommunications, University of West London | 2019-12-05 |
| OpenAI Five | Games | 6.7e+22 | OpenAI | 2019-12-13 |
| OpenAI Five Rerun | Games | 1.3e+22 | OpenAI | 2019-12-13 |
| DD-PPO | Robotics | 7.8e+20 | Georgia Institute of Technology, Facebook AI Research, Oregon State University, Simon Fraser University | 2019-12-19 |
| AlphaFold | Biology | 1.0e+20 | DeepMind | 2020-01-15 |
| ContextNet + Noisy Student | Speech | 8.2e+21 | 2020-01-19 | |
| Meena | Language | 1.1e+23 | Google Brain | 2020-01-28 |
| TaLK Convolution | Language | 2.7e+19 | Carleton University | 2020-02-08 |
| ALBERT-xxlarge | Language | 2.4e+21 | Toyota Technological Institute at Chicago, Google | 2020-02-09 |
| Turing-NLG | Language | 1.6e+22 | Microsoft | 2020-02-13 |
| Feedback Transformer | Language | 7.7e+18 | LORIA, University of Lorraine, Facebook AI Research | 2020-02-21 |
| TransformerXL + spectrum control | Language | 2.6e+19 | University of California Los Angeles (UCLA), JD.com | 2020-03-11 |
| Tensor-Transformer(1core)+PN (WT103) | Language | 1.6e+18 | University of California (UC) Berkeley | 2020-03-17 |
| ELECTRA | Language | 3.1e+21 | Stanford University, Google, Google Brain | 2020-03-23 |
| MetNet | Earth science | 9.5e+18 | 2020-03-24 | |
| Once for All | Vision | 6.2e+20 | MIT-IBM Watson AI Lab, Massachusetts Institute of Technology (MIT), IBM | 2020-04-29 |
| UnifiedQA | Language | 1.6e+19 | Allen Institute for AI, University of Washington | 2020-05-02 |
| DETR | Vision | 4.0e+20 | 2020-05-26 | |
| GPT-3 175B (davinci) | Language | 3.1e+23 | OpenAI | 2020-05-28 |
| GShard (dense) | Language | 4.8e+22 | 2020-06-30 | |
| DeLighT | Language | 3.8e+18 | University of Washington, Allen Institute for AI, Facebook AI Research | 2020-08-03 |
| ERNIE-GEN (large) | Language | 2.0e+20 | Baidu | 2020-08-06 |
| ProBERTa | Biology | 9.7e+18 | University of Illinois Urbana-Champaign (UIUC), Reed College | 2020-09-01 |
| LUKE | Language | 1.8e+22 | University of Washington, National Institute of Informatics | 2020-10-02 |
| Conformer + Wav2vec 2.0 + Noisy Student | Speech | 7.6e+21 | Google, Google Research, Google Brain | 2020-10-20 |
| mT5-XXL | Language | 8.2e+22 | Google, Google Research | 2020-10-20 |
| German ELECTRA Large | Language | 1.4e+21 | deepset, Bayerische Staatsbibliothek Muenchen | 2020-10-21 |
| ViT-Huge/14 | Vision | 4.3e+21 | Google Brain, Google Research | 2020-10-22 |
| wave2vec 2.0 LARGE | Speech | 3.9e+21 | 2020-10-22 | |
| KEPLER | Language | 1.7e+21 | Tsinghua University, Mila - Quebec AI (originally Montreal Institute for Learning Algorithms), HEC, CIFAR AI Research, Princeton University, University of Montreal / Université de Montréal | 2020-11-23 |
| AlphaFold 2 | Biology | 3.0e+21 | DeepMind | 2020-11-30 |
| CPM-Large | Language | 2.6e+20 | Tsinghua University, Beijing Academy of Artificial Intelligence / BAAI | 2020-12-01 |
| ESM1b | Biology | 5.1e+21 | Facebook AI Research, New York University (NYU) | 2020-12-15 |
| DensePhrases | Language | 2.1e+18 | Korea University, Princeton University | 2020-12-23 |
| CT-MoS (WT2) | Language | 5.4e+17 | Google, National Tsing Hua University | 2020-12-25 |
| ERNIE-Doc (247M) | Language | 3.0e+19 | Baidu | 2020-12-31 |
| CLIP (ViT L/14@336px) | Multimodal, Vision, Language, Video | 1.0e+22 | OpenAI | 2021-01-05 |
| DALL-E | Image generation | 4.7e+22 | OpenAI | 2021-01-05 |
| Switch | Language | 8.2e+22 | 2021-01-11 | |
| DeiT-B | Vision | 7.9e+19 | Meta AI, Sorbonne University | 2021-01-15 |
| DLWP | Earth science | 5.7e+18 | University of Washington, Microsoft Research | 2021-02-09 |
| MSA Transformer | Biology | 5.5e+21 | Facebook AI Research, University of California (UC) Berkeley, New York University (NYU) | 2021-02-13 |
| SRU++ Large | Language | 2.1e+19 | ASAPP | 2021-02-24 |
| Meta Pseudo Labels | Vision | 4.8e+22 | Google Brain, Google AI | 2021-03-01 |
| Generative BST | Language | 1.4e+22 | Facebook AI Research | 2021-03-05 |
| M6-T | Multimodal, Language, Vision | 5.5e+21 | Alibaba | 2021-03-05 |
| PLUG | Language | 3.6e+22 | Alibaba | 2021-04-19 |
| ProtT5-XL-U50 | Biology | 1.9e+22 | Technical University of Munich, Med AI Technology, NVIDIA, Oak Ridge National Laboratory, Google, Seoul National University | 2021-05-04 |
| ProtBERT-BFD | Biology | 3.9e+22 | Technical University of Munich, NVIDIA, Seoul National University, Google, Oak Ridge National Laboratory, Med AI Technology | 2021-05-04 |
| ADM | Image generation | 6.2e+21 | OpenAI | 2021-05-11 |
| MedBERT | Medicine | 9.5e+18 | Peng Cheng Laboratory, University of Texas at Houston | 2021-05-20 |
| Transformer local-attention (NesT-B) | Vision | 2.4e+19 | Google Cloud, Google Research | 2021-05-26 |
| CogView | Image generation | 2.7e+22 | Tsinghua University, Alibaba DAMO Academy | 2021-05-26 |
| ByT5-XXL | Language | 8.1e+22 | Google, Google Research | 2021-05-28 |
| ViT-G/14 | Vision | 5.8e+22 | Google Brain, Google Research | 2021-06-08 |
| CoAtNet | Vision | 4.3e+22 | Google, Google Research, Google Brain | 2021-06-09 |
| EMDR | Language | 1.9e+21 | Mila - Quebec AI (originally Montreal Institute for Learning Algorithms), McGill University, DeepMind | 2021-06-09 |
| DeBERTa | Language | 2.6e+22 | Microsoft | 2021-06-10 |
| ALIGN | Multimodal, Vision, Language | 2.6e+22 | Google Research | 2021-06-11 |
Denoising Diffusion Probabilistic Models (LSUN Bedroom) | Vision | 7.8e+19 | University of California (UC) Berkeley | 2021-06-11 |
| StyleGAN3-R | Image generation | 2.4e+21 | NVIDIA, Aalto University | 2021-06-21 |
| StyleGAN3-T | Image generation | 1.7e+21 | NVIDIA, Aalto University | 2021-06-21 |
| EfficientNetV2-XL | Vision | 9.6e+19 | Google, Google Brain | 2021-06-23 |
| Fold2Seq | Biology | 1.4e+17 | IBM, Texas A&M | 2021-06-24 |
| Adaptive Input Transformer + RD | Language | 8.6e+19 | Microsoft Research Asia, Soochow University | 2021-06-28 |
| ERNIE 3.0 | Language | 2.2e+22 | Baidu | 2021-07-05 |
| Codex | Language | 7.3e+22 | OpenAI | 2021-07-07 |
| GOAT | Games | 2.4e+22 | DeepMind | 2021-07-27 |
| HuBERT | Speech | 5.5e+21 | Facebook AI Research | 2021-07-27 |
| SEER | Vision | 1.8e+22 | Facebook AI Research, INRIA | 2021-07-29 |
| YOLOX-X | Vision | 6.3e+20 | Megvii Inc | 2021-08-06 |
| Jurassic-1-Jumbo | Language | 3.7e+23 | AI21 Labs | 2021-08-11 |
| Zidong Taichu | Multimodal, Speech, Vision, Language | 8.0e+20 | Chinese Academy of Sciences, Wuhan AI Computing Center | 2021-08-11 |
| DNABERT | Biology | 1.1e+20 | Northeastern University | 2021-08-15 |
| XLMR-XXL | Language | 3.4e+22 | Facebook AI Research | 2021-08-17 |
| FLAN 137B | Language | 2.0e+24 | Google Research | 2021-09-03 |
| PermuteFormer | Language | 2.8e+18 | Peking University | 2021-09-06 |
| HyperCLOVA 204B | Language | 2.0e+23 | NAVER | 2021-09-10 |
| PLATO-XL | Language | 9.9e+21 | Baidu | 2021-09-20 |
| Turing ULRv5 | Language | 2.9e+22 | Microsoft | 2021-09-28 |
| AlphaFold-Multimer | Biology | 4.4e+21 | Google DeepMind, DeepMind | 2021-10-04 |
| Megatron-Turing NLG 530B | Language | 8.6e+23 | Microsoft, NVIDIA | 2021-10-11 |
| Yuan 1.0 | Language | 3.5e+23 | Inspur | 2021-10-12 |
| base LM+GNN+kNN | Language | 5.3e+19 | Shannon.AI, Nanjing University, Nanyang Technological University, Zhejiang University (ZJU) | 2021-10-17 |
| S4 | Language | 7.8e+19 | Stanford University | 2021-10-31 |
| CodeT5-base | Language | 1.6e+21 | Salesforce, Nanyang Technological University | 2021-11-01 |
| Projected GAN | Image generation | 1.0e+19 | Heidelberg University | 2021-11-01 |
| Masked Autoencoders ViT-H | Vision | 4.6e+20 | Facebook AI Research | 2021-11-11 |
| Swin Transformer V2 (SwinV2-G) | Vision, Video | 1.1e+21 | Microsoft Research Asia | 2021-11-18 |
| BASIC-L | Vision | 4.1e+22 | 2021-11-19 | |
| Florence | Vision | 4.8e+22 | Microsoft | 2021-11-22 |
| NÜWA | Multimodal, Vision, Image generation, Video, Language | 7.2e+21 | Microsoft Research, Peking University | 2021-11-24 |
| Student of Games | Games | 3.7e+22 | DeepMind | 2021-12-06 |
| Gopher (280B) | Language | 6.3e+23 | DeepMind | 2021-12-08 |
| GLaM | Language | 3.6e+23 | 2021-12-13 | |
| Contriever | Language | 1.6e+20 | Meta AI, University College London (UCL), PSL University, Université Grenoble Alpes | 2021-12-16 |
| XGLM-7.5B | Language | 2.2e+22 | Meta AI, Facebook AI Research | 2021-12-20 |
| ERNIE 3.0 Titan | Language | 1.0e+24 | Baidu, Peng Cheng Laboratory | 2021-12-23 |
| Detic | Vision | 2.3e+19 | Meta AI, University of Texas at Austin | 2022-01-07 |
| InstructGPT 175B | Language | 3.2e+23 | OpenAI | 2022-01-27 |
| AlphaCode | Language | 2.4e+23 | DeepMind | 2022-02-02 |
| RETRO-7B | Language | 1.7e+22 | DeepMind | 2022-02-07 |
| GPT-NeoX-20B | Language | 9.3e+22 | EleutherAI | 2022-02-09 |
| LaMDA | Language | 3.6e+23 | 2022-02-10 | |
| ProteinBERT | Biology | 6.5e+19 | Hebrew University of Jerusalem, Ben-Gurion University of the Negev, Deep Trading | 2022-02-10 |
| ST-MoE | Language | 2.9e+23 | Google, Google Brain, Google Research | 2022-02-17 |
| FourCastNet | Earth science | 3.5e+20 | NVIDIA, NERSC, Lawrence Berkeley National Laboratory, University of Michigan, Rice University, California Institute of Technology, Purdue University | 2022-02-22 |
| PolyCoder | Language | 1.1e+21 | Carnegie Mellon University (CMU) | 2022-02-26 |
| ViT-G (model soup) | Vision | 3.4e+21 | University of Washington, Columbia University, Google, Meta AI, Tel Aviv University | 2022-03-10 |
| Segatron-XL large, M=384 + HCP | Language | 2.6e+19 | Microsoft Research, University of Waterloo | 2022-03-21 |
| Make-A-Scene | Image generation | 6.4e+21 | Meta AI | 2022-03-24 |
| Chinchilla | Language | 5.8e+23 | DeepMind | 2022-03-29 |
| PaLM (540B) | Language | 2.5e+24 | Google Research | 2022-04-04 |
| DALL·E 2 | Image generation | 3.4e+23 | OpenAI | 2022-04-06 |
| BERT-RBP | Biology | 1.4e+20 | Waseda University | 2022-04-07 |
| Stable Diffusion (LDM-KL-8-G) | Image generation | 5.0e+22 | Runway, Ludwig Maximilian University of Munich, Heidelberg University | 2022-04-13 |
| Sparse all-MLP | Language | 5.3e+20 | Meta AI | 2022-04-14 |
| Flamingo | Multimodal, Vision, Language, Video | 2.2e+23 | DeepMind | 2022-04-29 |
| OPT-175B | Language | 4.3e+23 | Meta AI | 2022-05-02 |
| UL2 | Language | 1.2e+23 | Google Research, Google Brain | 2022-05-10 |
| Gato | Multimodal, Robotics, Games, Language | 4.0e+21 | DeepMind | 2022-05-12 |
| Imagen | Image generation | 1.5e+22 | Google Brain | 2022-05-23 |
| GPT-2 Medium (FlashAttention) | Language | 8.9e+20 | Stanford University, University at Buffalo | 2022-05-27 |
| Tranception | Biology | 7.2e+21 | University of Oxford, Harvard Medical School, Cohere | 2022-05-27 |
| DITTO | Language | 3.3e+18 | Tsinghua University, Apple, Westlake University, Chinese University of Hong Kong (CUHK) | 2022-06-06 |
| CoCa | Vision | 7.3e+22 | Google Research | 2022-06-14 |
| Parti | Image generation | 5.1e+23 | Google Research | 2022-06-22 |
| ProGen2-xlarge | Biology | 1.4e+22 | Salesforce Research, Columbia University, Johns Hopkins University | 2022-06-27 |
| Minerva (540B) | Language | 2.7e+24 | 2022-06-29 | |
| CodeT5-large | Language | 2.7e+21 | Salesforce | 2022-07-05 |
| NLLB | Language | 1.8e+22 | Meta AI | 2022-07-06 |
| BLOOM-176B | Language | 3.7e+23 | Hugging Face, BigScience | 2022-07-11 |
| ESM2-15B | Biology | 7.4e+22 | Meta AI, New York University (NYU), Stanford University, Massachusetts Institute of Technology (MIT) | 2022-07-21 |
| OmegaPLM | Biology | 1.0e+22 | Massachusetts Institute of Technology (MIT), Westlake University | 2022-07-22 |
| AlexaTM 20B | Language | 2.0e+23 | Amazon | 2022-08-02 |
| GLM-130B | Language | 3.5e+23 | Tsinghua University | 2022-08-04 |
| BlenderBot 3 | Language | 4.3e+23 | McGill University, Meta AI, Mila - Quebec AI (originally Montreal Institute for Learning Algorithms) | 2022-08-10 |
| BEIT-3 | Multimodal, Vision, Language | 7.0e+19 | Microsoft | 2022-08-22 |
| PaLI | Language, Vision, Multimodal | 1.7e+23 | 2022-09-14 | |
| Whisper | Speech | 4.2e+21 | OpenAI | 2022-09-21 |
| DiffDock | Biology | 7.2e+19 | Massachusetts Institute of Technology (MIT) | 2022-10-04 |
| AlphaTensor | Other, Games, Mathematics | 7.1e+20 | DeepMind | 2022-10-05 |
| GenSLM | Biology | 1.4e+21 | University of Chicago, NVIDIA, Harvard University, Cerebras Systems, Technical University of Munich, California Institute of Technology | 2022-10-11 |
| U-PaLM (540B) | Language | 2.5e+24 | 2022-10-20 | |
| Flan-PaLM 540B | Language | 2.5e+24 | 2022-10-20 | |
| eDiff-I | Image generation | 5.5e+19 | NVIDIA | 2022-11-02 |
| Mogrifier RLSTM (WT2) | Language | 1.4e+17 | DeepMind | 2022-11-03 |
| InternImage | Vision | 2.4e+21 | Shanghai AI Lab, Tsinghua University, Nanjing University, SenseTime, Chinese University of Hong Kong (CUHK) | 2022-11-10 |
| EVA-01 | Vision | 1.5e+22 | Beijing Academy of Artificial Intelligence / BAAI, Huazhong University of Science and Technology, Zhejiang University (ZJU), Beijing Institute of Technology | 2022-11-14 |
| Galactica | Language, Biology | 3.2e+23 | Meta AI | 2022-11-16 |
| Fusion in Encoder | Language | 1.3e+20 | Samsung | 2022-11-18 |
| AR-LDM | Image generation | 5.1e+20 | Alibaba, University of Waterloo, Vector Institute | 2022-11-20 |
| Discriminator Guidance | Image generation | 2.2e+20 | Korea Advanced Institute of Science and Technology (KAIST), NAVER | 2022-11-28 |
| GPT-3.5 | Language | 2.6e+24 | OpenAI | 2022-11-28 |
| Vega v2 | Language | 7.8e+22 | Wuhan University, JD Explore Academy, Shanghai AI Lab, Nanyang Technological University, Washington University in St Louis, Chongqing University of Posts and Telecommunications, University of Sydney | 2022-12-04 |
| CaLM | Biology | 2.9e+19 | University of Oxford | 2022-12-19 |
| Hybrid H3-2.7B | Language | 6.5e+21 | Stanford University, University at Buffalo | 2022-12-28 |
| VALL-E | Audio, Speech | 1.0e+19 | Microsoft | 2023-01-05 |
| DreamerV3 | Multimodal, Games | 2.2e+20 | DeepMind, University of Toronto | 2023-01-10 |
| Nucleotide Transformer | Biology | 8.1e+21 | NVIDIA, Technical University of Munich, InstaDeep | 2023-01-15 |
| Ankh_large | Biology | 6.5e+21 | Technical University of Munich, Columbia University | 2023-01-16 |
| DDPM-IP (CelebA) | Image generation | 3.5e+20 | Utrecht University | 2023-01-27 |
| BLIP-2 (Q-Former) | Vision, Language | 1.2e+21 | Salesforce Research | 2023-01-30 |
| ViT-22B | Vision | 1.9e+23 | 2023-02-10 | |
| LLaMA-65B | Language | 5.5e+23 | Meta AI | 2023-02-24 |
| DiT-XL/2 | Image generation | 6.0e+20 | New York University (NYU), University of California (UC) Berkeley | 2023-03-02 |
| AudioGen | Audio | 9.5e+21 | Meta AI, Hebrew University of Jerusalem | 2023-03-05 |
| Falcon-40B | Language | 2.4e+23 | Technology Innovation Institute | 2023-03-15 |
| GPT-4 | Multimodal, Language, Vision | 2.1e+25 | OpenAI | 2023-03-15 |
| PanGu-Σ | Language | 4.7e+23 | Huawei Noah's Ark Lab | 2023-03-20 |
| SigLIP 400M | Vision | 4.9e+21 | Google DeepMind | 2023-03-27 |
| VideoMAE V2 | Video | 9.7e+21 | Nanjing University, Shenzhen Institute of Advanced Technology, Shanghai AI Lab | 2023-03-29 |
| BloombergGPT | Language | 2.4e+23 | Bloomberg, Johns Hopkins University | 2023-03-30 |
| Segment Anything Model | Vision | 7.8e+21 | Meta AI | 2023-04-05 |
| Incoder-6.7B | Language | 3.0e+21 | Facebook AI Research, University of Washington, University of California (UC) Berkeley, Carnegie Mellon University (CMU), Toyota Technological Institute at Chicago | 2023-04-09 |
| DINOv2 | Vision | 7.4e+21 | Facebook AI Research, INRIA | 2023-04-14 |
| LLaVA | Multimodal, Vision, Language | 7.8e+22 | University of Wisconsin Madison, Microsoft Research, Columbia University | 2023-04-17 |
| StarCoder | Language | 8.5e+22 | Hugging Face, ServiceNow, Northeastern University, Mila - Quebec AI (originally Montreal Institute for Learning Algorithms), Carnegie Mellon University (CMU), Johns Hopkins University, Leipzig University, ScaDS.AI, Queen Mary University of London, Roblox, Sea AI Lab, Technion - Israel Institute of Technology, Monash University, CSIRO, Data61, McGill University, Saama, University of British Columbia (UBC), Massachusetts Institute of Technology (MIT), Technical University of Munich, IBM, University of Vermont, UnfoldML, SAP, University of Notre Dame, Columbia University, New York University (NYU), University of Allahabad, Discover Dollar, Toloka, Telefonica, Stanford University, Weizmann Institute of Science, Alan Turing Institute, Wellesley College, EleutherAI, Forschungszentrum Julich | 2023-05-09 |
| PaLM 2 | Language | 7.3e+24 | 2023-05-10 | |
| InstructBLIP | Multimodal, Language, Vision | 1.9e+20 | Salesforce Research, Hong Kong University of Science and Technology (HKUST), Nanyang Technological University | 2023-05-11 |
| ONE-PEACE | Multimodal, Vision, Speech, Language | 1.8e+20 | Alibaba, Huazhong University of Science and Technology | 2023-05-18 |
| PaLI-X | Multimodal, Language, Vision, Video | 5.6e+23 | Google Research | 2023-05-29 |
| HyenaDNA | Biology | 1.8e+21 | Stanford University, Harvard University, Mila - Quebec AI (originally Montreal Institute for Learning Algorithms), University of Montreal / Université de Montréal | 2023-06-27 |
| Pangu-Weather | Earth science | 4.0e+22 | Huawei | 2023-07-05 |
| InternLM | Language | 1.0e+24 | Shanghai AI Lab, SenseTime | 2023-07-06 |
| xTrimoPGLM -100B | Biology | 6.2e+23 | Tsinghua University, BioMap Research | 2023-07-06 |
| Claude 2 | Language | 3.9e+24 | Anthropic | 2023-07-11 |
| Llama 2-7B | Language | 8.4e+22 | Meta AI | 2023-07-18 |
| Llama 2-70B | Language | 8.1e+23 | Meta AI | 2023-07-18 |
| AudioLM | Audio | 3.9e+18 | Google Research | 2023-07-26 |
| GGNN | Biology | 7.6e+21 | Westlake University, Tsinghua University, Toyota Technological Institute at Chicago | 2023-08-05 |
| PeptideBERT | Biology | 4.9e+16 | Carnegie Mellon University (CMU) | 2023-08-28 |
| Jais | Language | 4.9e+22 | Cerebras Systems, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Inception G42 | 2023-08-29 |
| Swift | Robotics | 5.3e+16 | Intel Labs | 2023-08-30 |
| Falcon-180B | Language | 3.8e+24 | Technology Innovation Institute | 2023-09-06 |
| Amazon Titan | Language, Image generation | 4.8e+24 | Amazon | 2023-09-28 |
| FinGPT-13B | Language | 1.6e+23 | University of California Los Angeles (UCLA), Columbia University, New York University (NYU) | 2023-10-07 |
| RoseTTAFold All-Atom (RFAA) | Biology | 2.1e+20 | University of Washington, Seoul National University, University of Sheffield | 2023-10-09 |
| CODEFUSION (Python) | Language | 7.9e+18 | Microsoft, Microsoft Research | 2023-10-26 |
| ChatGLM3-6B | Multimodal, Language, Vision | 5.0e+22 | Zhipu AI | 2023-10-27 |
| Skywork-13B | Language | 2.5e+23 | Kunlun Inc. | 2023-10-30 |
| Yi-34B | Language | 6.1e+23 | 01.AI | 2023-11-02 |
| Grok-1 | Language | 2.9e+24 | xAI | 2023-11-04 |
| LLaVA 1.5 | Multimodal, Language, Vision | 7.8e+22 | University of Wisconsin Madison, Microsoft Research | 2023-11-05 |
| GPT-4 Turbo | Multimodal, Vision, Language, Image generation | 2.2e+25 | OpenAI | 2023-11-06 |
| CogVLM-17B | Multimodal, Vision, Language | 6.3e+22 | Tsinghua University, Zhipu AI, Beihang University | 2023-11-06 |
| MultiBand Diffusion | Audio, Speech | 2.6e+19 | Meta AI, Hebrew University of Jerusalem, LORIA | 2023-11-08 |
| Volcano 13B | Language, Multimodal, Vision | 4.6e+22 | Korea University, Korea Advanced Institute of Science and Technology (KAIST), LG | 2023-11-13 |
| SPHINX (Llama 2 13B) | Vision, Language, Multimodal | 3.0e+22 | Shanghai AI Lab, Chinese University of Hong Kong (CUHK), ShanghaiTech University | 2023-11-13 |
| GraphCast | Earth science | 2.1e+22 | Google DeepMind | 2023-11-14 |
| Nemotron-3-8B | Language | 1.8e+23 | NVIDIA | 2023-11-15 |
| Inflection-2 | Language | 1.0e+25 | Inflection AI | 2023-11-22 |
| Qwen-72B | Language | 1.3e+24 | Alibaba | 2023-11-30 |
| Gemini 1.0 Pro | Multimodal, Language, Vision | 1.8e+24 | Google DeepMind | 2023-12-06 |
| Gemini 1.0 Ultra | Multimodal, Language, Vision | 5.0e+25 | Google DeepMind | 2023-12-06 |
| Llama Guard | Language | 1.6e+23 | Meta AI | 2023-12-07 |
| Mixtral 8x7B | Language | 7.7e+23 | Mistral AI | 2023-12-11 |
| VILA-13B | Multimodal | 2.3e+21 | NVIDIA, Massachusetts Institute of Technology (MIT) | 2023-12-12 |
| CogAgent | Vision, Language | 6.7e+22 | Tsinghua University, Zhipu AI | 2023-12-14 |
| FunSearch | Language, Search | 3.9e+23 | Google DeepMind | 2023-12-14 |
| nekomata-14b | Language | 2.6e+23 | rinna | 2023-12-21 |
| Qwen1.5-72B | Language | 1.3e+24 | Alibaba | 2024-02-04 |
| Gemini 1.5 Pro | Language, Multimodal | 1.6e+25 | Google DeepMind | 2024-02-15 |
| Stable Diffusion 3 | Image generation | 5.0e+22 | Stability AI | 2024-02-22 |
| MegaScale (Production) | Language | 3.9e+24 | ByteDance, Peking University | 2024-02-23 |
| Mistral Large | Language | 1.1e+25 | Mistral AI | 2024-02-26 |
| Aramco Metabrain AI | Language | 1.0e+25 | Saudi Aramco | 2024-03-04 |
| Claude 3 Opus | Multimodal, Language, Vision | 1.6e+25 | Anthropic | 2024-03-04 |
| Inflection-2.5 | Language | 8.0e+24 | Inflection AI | 2024-03-07 |
| MM1-30B | Multimodal, Language, Vision | 4.9e+23 | Apple | 2024-03-14 |
| DBRX | Language | 2.6e+24 | Databricks | 2024-03-27 |
| Reka Core | Multimodal, Language, Vision, Video, Speech | 8.4e+24 | Reka AI | 2024-04-15 |
| Llama 3-70B | Language | 7.9e+24 | Meta AI | 2024-04-18 |
| GenCast | Earth science | 8.2e+20 | Google DeepMind | 2024-05-01 |
| VILA1.5-13B | Multimodal, Language, Vision | 2.3e+21 | NVIDIA, Massachusetts Institute of Technology (MIT) | 2024-05-03 |
| AlphaFold 3 | Biology | 4.1e+22 | Google DeepMind, Isomorphic Labs | 2024-05-08 |
| Yi-Large | Language | 1.8e+24 | 01.AI | 2024-05-13 |
| GPT-4o | Multimodal, Language, Audio, Speech, Vision | 3.8e+25 | OpenAI | 2024-05-13 |
| Octo-Base | Robotics | 5.8e+20 | University of California (UC) Berkeley, Stanford University, Carnegie Mellon University (CMU), DeepMind | 2024-05-20 |
| ALLaM 7B | Language | 9.0e+22 | Saudi Data and Artificial Intelligence Authority | 2024-05-21 |
| ALLaM adapted 70B | Language | 1.1e+24 | Saudi Data and Artificial Intelligence Authority | 2024-05-21 |
| Qwen2-72B | Language | 3.0e+24 | Alibaba | 2024-06-07 |
| Llama-3.1-Nemotron-70B-Instruct | Language | 7.9e+24 | NVIDIA, Meta AI | 2024-06-12 |
| OpenVLA | Robotics, Vision, Language | 1.1e+23 | Stanford University, University of California (UC) Berkeley, Toyota Research Institute, Google DeepMind, Massachusetts Institute of Technology (MIT), Physical Intelligence | 2024-06-13 |
| Nemotron-4 340B | Language | 1.8e+25 | NVIDIA | 2024-06-14 |
| DeepSeek-Coder-V2 236B | Language | 1.3e+24 | DeepSeek | 2024-06-17 |
| Claude 3.5 Sonnet | Multimodal, Language, Vision | 2.7e+25 | Anthropic | 2024-06-20 |
| ESM3 (98B) | Biology | 1.1e+24 | EvolutionaryScale, University of California (UC) Berkeley | 2024-06-25 |
| GPT-4o mini | Language, Multimodal, Vision | 7.4e+24 | OpenAI | 2024-07-18 |
| Llama 3.1-405B | Language | 3.8e+25 | Meta AI | 2024-07-23 |
| Mistral Large 2 | Language | 2.1e+25 | Mistral AI | 2024-07-24 |
| AFM-server | Language | 4.3e+24 | Apple | 2024-07-29 |
| AFM-on-device | Language | 4.5e+23 | Apple | 2024-07-29 |
| LLaVA-OV-72B | Multimodal, Vision | 3.0e+24 | ByteDance, Nanyang Technological University, Chinese University of Hong Kong (CUHK), Hong Kong University of Science and Technology (HKUST) | 2024-08-06 |
| Grok-2 | Language, Vision, Multimodal | 3.0e+25 | xAI | 2024-08-13 |
| GLM-4-Plus | Language | 3.6e+25 | Zhipu AI | 2024-08-29 |
| DeepSeek-V2.5 | Language | 1.8e+24 | DeepSeek | 2024-09-06 |
| Qwen2.5-32B | Language | 3.5e+24 | Alibaba | 2024-09-17 |
| Qwen2.5 Instruct (72B) | Language | 7.9e+24 | Alibaba | 2024-09-19 |
| Qwen2.5-72B | Language | 7.8e+24 | Alibaba | 2024-09-19 |
| Telechat2-115B | Language | 6.9e+24 | China Telecom | 2024-09-20 |
| Llama 3.2 11B | Multimodal, Vision, Language | 5.8e+23 | Meta AI | 2024-09-24 |
| Movie Gen Video | Video, Vision | 1.6e+24 | Meta AI | 2024-10-04 |
| RDT-1B | Robotics | 4.1e+22 | Tsinghua University | 2024-10-10 |
| CHAI-1 | Biology | 7.8e+21 | Chai discovery | 2024-10-15 |
| Yi-Lightning | Language | 1.5e+24 | 01.AI | 2024-10-18 |
| NVLM-X 72B | Vision, Language | 3.0e+24 | NVIDIA | 2024-10-22 |
| NVLM-H 72B | Vision, Language | 3.0e+24 | NVIDIA | 2024-10-22 |
| NVLM-D 72B | Vision, Language | 3.0e+24 | NVIDIA | 2024-10-22 |
| Doubao-pro | Language | 2.5e+25 | ByteDance | 2024-10-28 |
| Hunyuan-Large | Language | 3.5e+24 | Tencent | 2024-11-06 |
| Amazon Nova Pro | Multimodal, Language, Video, Vision | 6.0e+24 | Amazon | 2024-12-03 |
| Llama 3.3 70B | Language | 6.9e+24 | Meta AI | 2024-12-06 |
| EXAONE 3.5 32B | Language | 1.3e+24 | LG AI Research | 2024-12-09 |
| DeepSeek-V3 | Language | 3.4e+24 | DeepSeek | 2024-12-24 |
| DeepSeek-R1 | Language | 4.0e+24 | DeepSeek | 2025-01-20 |
| Eagle 2 | Vision, Robotics, Language | 4.7e+22 | NVIDIA, Nanjing University, Tsinghua University, Hong Kong Polytechnic University, Johns Hopkins University, New York University (NYU) | 2025-01-20 |
| Grok 3 | Language, Vision, Multimodal | 3.5e+26 | xAI | 2025-02-17 |
| Claude 3.7 Sonnet | Language, Vision, Multimodal | 3.3e+25 | Anthropic | 2025-02-24 |
| GPT-4.5 | Language, Vision, Multimodal | 2.1e+26 | OpenAI | 2025-02-27 |
| QwQ-32B | Language | 3.5e+24 | Alibaba | 2025-03-06 |
| EXAONE Deep 32B | Language | 1.3e+24 | LG AI Research | 2025-03-16 |
| Llama 4 Scout | Multimodal, Language, Vision | 4.1e+24 | Meta AI | 2025-04-05 |
| Llama 4 Maverick | Multimodal, Language, Vision | 2.2e+24 | Meta AI | 2025-04-05 |
| Llama 4 Behemoth (preview) | Multimodal, Language, Vision | 5.2e+25 | Meta AI | 2025-04-05 |
| Pangu Ultra | Language | 1.1e+25 | Huawei | 2025-04-10 |
| Qwen3-235B-A22B | Language | 4.8e+24 | Alibaba | 2025-04-29 |
| Seed1.5-VL | Vision, Language, Multimodal, Video | 1.4e+24 | ByteDance | 2025-05-11 |
| FGN | Earth science | 9.6e+21 | Google DeepMind | 2025-06-12 |
| Grok 4 | Language, Multimodal, Vision, Speech | 5.0e+26 | xAI | 2025-07-09 |
| Kimi K2 | Language | 3.0e+24 | Moonshot | 2025-07-11 |
| EXAONE 4.0 (32B) | Language | 2.7e+24 | LG AI Research | 2025-07-15 |
| Qwen3-Coder-480B-A35B | Language | 1.6e+24 | Alibaba | 2025-07-22 |
| GLM 4.5 | Language | 4.4e+24 | Zhipu AI, Tsinghua University | 2025-08-05 |
| gpt-oss-20b | Language | 5.5e+23 | OpenAI | 2025-08-05 |
| gpt-oss-120b | Language | 4.9e+24 | OpenAI | 2025-08-05 |
| Qwen3-Max | Language | 1.5e+25 | Alibaba | 2025-09-05 |
Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.
Learn more about this graph
Data come from Epoch AI’s AI Models database, which contains information on over 2700 models trained since 1950. We begin by filtering to models meeting one of our notability criteria, and to models trained after 2010, in order to focus on recent trends in AI. We additionally filter out any models missing values for either publication date or training compute, leaving us with a final dataset of 428 observations.
Analysis
We perform a simple log-linear regression to obtain a trend in training compute over time, using the slope coefficient’s standard errors to estimate a 90% confidence interval. For a more extensive justification of the log-linear trend, see Appendix 1: Compute-Trend Model Selection in our full report.
Results are presented in the table below.
| R2 | Annual growth | 90% confidence interval |
|---|---|---|
| 0.60 | 4.7x / year | (4.3x to 5.2x) |
Assumptions and limitations
- We present results for all notable models. Different subsets of the data present somewhat different trends (e.g. frontier models, language models, etc.)
- Our training compute figures are estimated with some uncertainty. We assume our estimates are unbiased, and that measurement noise does not substantially invalidate our findings.

