Training frontier models requires a large and growing amount of power for GPUs, servers, cooling and other equipment. This is driven by an increase in GPU count; power draw per GPU is also growing, but at only a few percent per year.
| Model | Domain | Training compute (FLOP) | Training power draw (W) | Organization | Publication date |
|---|---|---|---|---|---|
| Deep Autoencoders | Vision | 3.7e+16 | 3.9e+2 | University of Toronto | 2011-04-29 |
| High Performance CNN (NORB) | Vision | 2.6e+16 | 1.9e+3 | IDSIA,SUPSI | 2011-07-16 |
| CNN Committee (NIST) | Vision | 2.6e+16 | 1.9e+3 | IDSIA | 2011-09-18 |
| CNN Committee (MNIST) | Vision | 5.2e+16 | 1.9e+3 | IDSIA | 2011-09-18 |
| DNN EM segmentation | Vision | 4.8e+17 | 1.9e+3 | IDSIA,SUPSI | 2012-12-03 |
| VGG16 | Vision | 1.2e+19 | 1.9e+3 | University of Oxford | 2014-09-04 |
| DeepSpeech2 (English) | Speech | 2.6e+19 | 7.6e+3 | Baidu Research - Silicon Valley AI Lab | 2015-12-08 |
| 10 LSTMS + KN-5 (OPTIMAL WEIGHTS) | Language | 8.8e+19 | 1.5e+4 | Google Brain | 2016-02-11 |
| BIG LSTM+CNN INPUTS | Language | 1.0e+20 | 1.5e+4 | Google Brain | 2016-02-11 |
| Xception | Vision | 4.4e+20 | 3.4e+4 | 2016-10-07 | |
| PolyNet | Vision | 6.4e+19 | 1.5e+4 | Chinese University of Hong Kong (CUHK) | 2016-11-17 |
| MoE-Multi | Language | 9.4e+19 | 3.0e+4 | Jagiellonian University,Google Brain | 2017-01-23 |
| JFT | Vision | 8.4e+20 | 2.9e+4 | Google Research,Carnegie Mellon University (CMU) | 2017-07-10 |
| AmoebaNet-A (F=448) | Vision | 3.9e+20 | 2.1e+5 | Google Brain | 2018-02-05 |
| ResNeXt-101 32x48d | Vision | 8.7e+21 | 1.9e+5 | 2018-05-02 | |
| Big Transformer for Back-Translation | Language | 4.8e+20 | 6.1e+4 | Facebook AI Research,Google Brain | 2018-08-28 |
| BigGAN-deep 512x512 | Image generation | 1.8e+21 | 1.1e+5 | Heriot-Watt University,DeepMind | 2018-09-28 |
| Sparse Transformer (ImageNet) | Image generation | 1.5e+21 | 3.7e+4 | OpenAI | 2019-04-23 |
| MnasNet-A1 + SSDLite | Vision | 1.5e+21 | 1.1e+5 | 2019-05-29 | |
| Grover-Mega | Language | 4.6e+21 | 5.4e+4 | University of Washington | 2019-05-29 |
| MnasNet-A3 | Vision | 1.5e+21 | 1.1e+5 | 2019-05-29 | |
| RoBERTa Large | Language | 8.5e+21 | 4.9e+5 | Facebook,University of Washington | 2019-07-01 |
| Megatron-LM (1.2B) | Language | 1.1e+22 | 5.7e+2 | NVIDIA | 2019-09-17 |
| Megatron-LM (8.3B) | Language | 9.1e+21 | 2.4e+5 | NVIDIA | 2019-09-17 |
| Megatron-BERT | Language | 2.2e+22 | 2.4e+5 | NVIDIA | 2019-09-17 |
| T5-11B | Language | 3.3e+22 | 2.1e+5 | 2019-10-23 | |
| AlphaStar | Games | 1.1e+23 | 1.6e+5 | DeepMind | 2019-10-30 |
| XLM-RoBERTa | Language | 2.1e+22 | 2.4e+5 | Facebook AI | 2019-11-05 |
| Noisy Student (L2) | Vision | 2.6e+22 | 4.3e+5 | Carnegie Mellon University (CMU),Google | 2019-11-11 |
| OpenAI Five Rerun | Games | 1.3e+22 | 2.4e+5 | OpenAI | 2019-12-13 |
| OpenAI Five | Games | 6.7e+22 | 7.3e+5 | OpenAI | 2019-12-13 |
| Meena | Language | 1.1e+23 | 4.3e+5 | Google Brain | 2020-01-28 |
| Turing-NLG | Language | 1.6e+22 | 1.2e+5 | Microsoft | 2020-02-13 |
| GPT-3 175B (davinci) | Language | 3.1e+23 | 4.8e+6 | OpenAI | 2020-05-28 |
| GShard (dense) | Language | 4.8e+22 | 4.3e+5 | 2020-06-30 | |
| DALL-E | Image generation | 4.7e+22 | 4.9e+5 | OpenAI | 2021-01-05 |
| Switch | Language | 8.2e+22 | 4.3e+5 | 2021-01-11 | |
| Meta Pseudo Labels | Vision | 4.8e+22 | 4.3e+5 | Google Brain,Google AI | 2021-03-01 |
| PLUG | Language | 3.6e+22 | 9.7e+4 | Alibaba | 2021-04-19 |
| PanGu-α | Language | 5.1e+22 | 1.2e+6 | Huawei Noah's Ark Lab | 2021-04-25 |
| ProtT5-XXL | Biology | 7.4e+22 | 2.1e+5 | Technical University of Munich,Med AI Technology,NVIDIA,Oak Ridge National Laboratory,Google,Seoul National University | 2021-05-04 |
| ViT-G/14 | Vision | 5.8e+22 | 8.6e+5 | Google Brain,Google Research | 2021-06-08 |
| Megatron-Turing NLG 530B | Language | 8.6e+23 | 3.4e+6 | Microsoft,NVIDIA | 2021-10-11 |
| Gopher (280B) | Language | 6.3e+23 | 1.7e+6 | DeepMind | 2021-12-08 |
| GLaM | Language | 3.6e+23 | 3.3e+5 | 2021-12-13 | |
| ERNIE 3.0 Titan | Language | 1.0e+24 | 1.1e+6 | Baidu,Peng Cheng Laboratory | 2021-12-23 |
| AlphaCode | Language | 2.4e+23 | 1.2e+6 | DeepMind | 2022-02-02 |
| LaMDA | Language | 3.6e+23 | 4.3e+5 | 2022-02-10 | |
| PaLM (540B) | Language | 2.5e+24 | 2.0e+6 | Google Research | 2022-04-04 |
| OPT-175B | Language | 4.3e+23 | 7.8e+5 | Meta AI | 2022-05-02 |
| LLaMA-65B | Language | 5.5e+23 | 1.6e+6 | Meta AI | 2023-02-24 |
| GPT-4 | Multimodal,Language,Vision | 2.1e+25 | 1.9e+7 | OpenAI | 2023-03-15 |
| Falcon-180B | Language | 3.8e+24 | 3.1e+6 | Technology Innovation Institute | 2023-09-06 |
| Amazon Titan | Language,Image generation | 4.8e+24 | 1.0e+7 | Amazon | 2023-09-28 |
| Inflection-2 | Language | 1.0e+25 | 6.7e+6 | Inflection AI | 2023-11-22 |
| Gemini 1.0 Ultra | Multimodal,Language,Vision | 5.0e+25 | 1.8e+7 | Google DeepMind | 2023-12-06 |
| MegaScale (Production) | Language | 3.9e+24 | 9.4e+6 | ByteDance,Peking University | 2024-02-23 |
| Nemotron-4 340B | Language | 1.8e+25 | 8.2e+6 | NVIDIA | 2024-06-14 |
| Llama 3.1-405B | Language | 3.8e+25 | 2.2e+7 | Meta AI | 2024-07-23 |
| Grok 3 | Language,Vision,Multimodal | 3.5e+26 | 1.1e+8 | xAI | 2025-02-17 |
| Llama 4 Behemoth (preview) | Multimodal,Language,Vision | 5.2e+25 | 4.3e+7 | Meta AI | 2025-04-05 |
| CNN committee (traffic sign) | Vision | 9.9e+14 | 1.9e+3 | IDSIA | 2011-10-03 |
| Large regularized LSTM | Language | 4.3e+16 | 4.3e+2 | New York University (NYU),Google Brain | 2014-09-08 |
| SPN-4+KN5 | Language | 4.4e+16 | 4.7e+2 | Singapore University of Technology & Design,DSO National Laboratories | 2014-09-14 |
| TC-DNN-BLSTM-DNN | Speech | 1.9e+17 | 4.3e+2 | Carnegie Mellon University (CMU) | 2015-04-06 |
| DCNN | Vision | 4.8e+17 | 4.7e+2 | University of Maryland,Rutgers University | 2015-08-07 |
| SAF R-CNN | Vision | 1.2e+19 | 4.8e+2 | Beijing Institute of Technology,Sun Yat-sen University,Panasonic R&D,National University of Singapore | 2015-10-28 |
| Named Entity Recognition model | Language | 9.7e+16 | 4.8e+2 | Carnegie Mellon University (CMU) | 2016-03-04 |
| Part-of-sentence tagging model | Language | 1.5e+17 | 4.8e+2 | Carnegie Mellon University (CMU) | 2016-05-29 |
| BIDAF | Language | 3.5e+18 | 3.8e+3 | University of Washington,Allen Institute for AI | 2016-11-05 |
| EnhanceNet | Vision | 1.3e+17 | 4.7e+2 | Max Planck Institute for Intelligent Systems | 2016-12-23 |
| VDCNN (on Amazon Review Full dataset) | Language | 5.7e+17 | 4.7e+2 | Facebook AI Research,University of Le Mans | 2017-01-27 |
| Transformer | Language | 7.4e+18 | 3.8e+3 | Google Research,Google Brain | 2017-06-12 |
| DeepLoc | Biology | 5.8e+17 | 4.8e+2 | Technical University of Denmark,University of Copenhagen | 2017-07-07 |
| RetinaNet-R101 | Vision | 2.1e+18 | 3.8e+3 | Facebook AI Research | 2017-08-07 |
| AlphaZero | Games | 1.1e+20 | 2.7e+6 | DeepMind | 2017-12-05 |
| Refined Part Pooling | Vision | 2.6e+16 | 9.5e+2 | Tsinghua University,University of Technology Sydney,University of Texas at San Antonio | 2018-01-09 |
| IMPALA | Games | 1.7e+20 | 4.8e+2 | DeepMind | 2018-02-05 |
| 4 layer QRNN (h=2500) | Language | 5.9e+17 | 4.5e+2 | Salesforce Research | 2018-03-22 |
| LSTM (Hebbian, Cache, MbPA) | Language | 3.3e+19 | 3.8e+3 | DeepMind,University College London (UCL) | 2018-03-27 |
| RNMT+ | Language | 1.8e+19 | 1.5e+4 | Google AI | 2018-04-26 |
| GPT-1 | Language | 1.8e+19 | 6.1e+2 | OpenAI | 2018-06-01 |
| DARTS | Language | 3.2e+17 | 4.8e+2 | DeepMind,Carnegie Mellon University (CMU) | 2018-06-24 |
Dexterous In-Hand Manipulation [control policy] | Robotics | 2.2e+20 | 4.6e+3 | OpenAI | 2018-08-01 |
| Transformer + Simple Recurrent Unit | Language | 1.1e+19 | 4.6e+3 | ASAPP,Cornell University,Google,Princeton University | 2018-09-17 |
| Transformer (Adaptive Input Embeddings) WT103 | Language | 4.5e+19 | 4.6e+3 | Facebook AI Research | 2018-09-28 |
| BERT-Large | Language | 2.8e+20 | 3.4e+4 | 2018-10-11 | |
Mesh-TensorFlow Transformer 2.9B (translation) | Language | 6.8e+19 | 3.4e+4 | Google Brain | 2018-11-05 |
| Mesh-TensorFlow Transformer 4.9B (language) | Language | 1.6e+20 | 1.4e+5 | Google Brain | 2018-11-05 |
| StyleGAN | Image generation | 3.9e+16 | 4.6e+3 | NVIDIA | 2018-12-12 |
| Transformer-XL (257M) | Language | 3.8e+20 | 1.3e+4 | Carnegie Mellon University (CMU),Google Brain | 2019-01-09 |
| Mono3D++ | 3D modeling,Vision | 4.9e+18 | 1.9e+3 | University of California Los Angeles (UCLA),Megvii Inc | 2019-01-11 |
| SSA | Biology | 1.3e+19 | 5.7e+2 | Massachusetts Institute of Technology (MIT) | 2019-02-22 |
| UniRep | Biology | 2.2e+19 | 2.3e+3 | Harvard University | 2019-03-26 |
| SciBERT | Language | 8.9e+19 | 1.7e+3 | Allen Institute for AI | 2019-03-26 |
| WeNet (Penn Treebank) | Language | 7.3e+17 | 5.7e+2 | Amazon | 2019-04-08 |
| MEGNet (molecule model) | Materials science | 4.5e+17 | 4.8e+2 | University of California San Diego | 2019-04-10 |
| MEGNet (crystal band gap model) | Materials science | 4.5e+17 | 4.8e+2 | University of California San Diego | 2019-04-10 |
| Sparse Transformer (Enwik8) | Language | 4.1e+18 | 4.6e+3 | OpenAI | 2019-04-23 |
| TAPE Transformer | Biology | 3.0e+19 | 2.3e+3 | University of California (UC) Berkeley,Covariant,Google,Chan Zuckerberg Initiative | 2019-06-19 |
| Tensorized Transformer (257M) | Language | 4.8e+18 | 9.5e+2 | Tianjin University,Microsoft Research Asia,Beijing Institute of Technology | 2019-06-24 |
| All-attention network + adaptive span | Language | 1.2e+19 | 3.7e+4 | Facebook AI Research | 2019-07-02 |
| trRosetta | Biology | 3.8e+19 | 5.3e+2 | Nankai University,University of Washington,Tianjin University,Harvard University | 2019-08-22 |
| Long-range sequence Compressive Transformers | Language | 1.0e+20 | 2.7e+4 | DeepMind | 2019-11-13 |
| FastSpeech | Speech | 7.2e+18 | 2.3e+3 | Zhejiang University (ZJU),Microsoft Research | 2019-11-20 |
| SeqVec | Biology | 4.1e+19 | 2.4e+3 | Technical University of Munich | 2019-12-17 |
| DD-PPO | Robotics | 7.8e+20 | 3.7e+4 | Georgia Institute of Technology,Facebook AI Research,Oregon State University,Simon Fraser University | 2019-12-19 |
| TaLK Convolution | Language | 2.7e+19 | 3.8e+3 | Carleton University | 2020-02-08 |
| ALBERT-xxlarge | Language | 2.4e+21 | 2.1e+5 | Toyota Technological Institute at Chicago,Google | 2020-02-09 |
| LSTM-3-layer+Gadam | Language | 2.6e+16 | 4.8e+2 | University of Oxford,University of Bristol,University of Cambridge | 2020-03-02 |
| TransformerXL + spectrum control | Language | 2.6e+19 | 2.3e+3 | University of California Los Angeles (UCLA),JD.com | 2020-03-11 |
| WDC20 / DLWP | Earth science | 2.4e+18 | 4.8e+2 | University of Washington,Microsoft Research | 2020-03-15 |
| MetNet | Earth science | 9.5e+18 | 1.1e+5 | 2020-03-24 | |
| AraBERT LArge v2 | Language | 1.5e+21 | 5.4e+4 | American University of Beirut | 2020-03-30 |
| AraBERT | Language | 3.2e+19 | 4.3e+3 | American University of Beirut | 2020-03-30 |
| Cube-Space AutoEncoder | Vision,Search | 1.1e+17 | 5.7e+2 | MIT-IBM Watson AI Lab | 2020-04-27 |
| UnifiedQA | Language | 1.6e+19 | 3.4e+3 | Allen Institute for AI,University of Washington | 2020-05-02 |
| NAS+ESS (156M) | Language | 2.9e+18 | 4.8e+2 | Northeastern University (China),Chinese Academy of Sciences,NiuTrans Research,Kingsoft | 2020-05-06 |
| mBART-50 | Language | 1.5e+22 | 1.5e+5 | Facebook AI | 2020-08-02 |
| DeLighT | Language | 3.8e+18 | 4.6e+3 | University of Washington,Allen Institute for AI,Facebook AI Research | 2020-08-03 |
| ProBERTa | Biology | 9.7e+18 | 2.3e+3 | University of Illinois Urbana-Champaign (UIUC),Reed College | 2020-09-01 |
| LUKE | Language | 1.8e+22 | 9.1e+3 | University of Washington,National Institute of Informatics | 2020-10-02 |
| Memformer (4 encoder + 16 decoder) | Language | 1.2e+19 | 1.9e+3 | UC Davis,Westlake University,Facebook AI | 2020-10-14 |
| Conformer + Wav2vec 2.0 + Noisy Student | Speech | 7.6e+21 | 1.1e+5 | Google,Google Research,Google Brain | 2020-10-20 |
| German ELECTRA Large | Language | 1.4e+21 | 2.7e+4 | deepset,Bayerische Staatsbibliothek Muenchen | 2020-10-21 |
| GBERT-Large | Language | 2.2e+21 | 2.7e+4 | deepset,Bayerische Staatsbibliothek Muenchen | 2020-10-21 |
| ChemBERTa | Biology | 8.5e+18 | 5.7e+2 | University of Toronto,Reverie Labs,DeepChem | 2020-10-23 |
| CPM-Large | Language | 2.6e+20 | 3.7e+4 | Tsinghua University,Beijing Academy of Artificial Intelligence / BAAI | 2020-12-01 |
| Profile Prediction | Biology | 5.0e+20 | 4.6e+3 | University of Washington,Salesforce Research | 2020-12-01 |
| RoBERTa (PFAM) | Biology | 1.2e+19 | 1.9e+3 | IBM Research,ETH Zurich | 2020-12-05 |
| ESM1b | Biology | 5.1e+21 | 7.3e+4 | Facebook AI Research,New York University (NYU) | 2020-12-15 |
| DensePhrases | Language | 2.1e+18 | 3.8e+3 | Korea University,Princeton University | 2020-12-23 |
| CT-MoS (WT2) | Language | 5.4e+17 | 1.9e+3 | Google,National Tsing Hua University | 2020-12-25 |
| AraGPT2-Mega | Language | 2.0e+21 | 5.4e+4 | American University of Beirut | 2020-12-31 |
| Subformer (122M) | Language | 5.1e+18 | 4.6e+3 | National Institute of Advanced Industrial Science and Technology (AIST),The University of Tokyo | 2021-01-01 |
| CLIP (ViT L/14@336px) | Multimodal,Vision,Language,Video | 1.0e+22 | 1.5e+5 | OpenAI | 2021-01-05 |
| NVAE (FFHQ) | Image generation | 5.2e+20 | 1.1e+4 | NVIDIA | 2021-01-08 |
| NVAE (Celeba HQ) | Image generation | 3.0e+20 | 1.1e+4 | NVIDIA | 2021-01-08 |
| NVAE (CIFAR 10) | Image generation | 5.9e+19 | 3.8e+3 | NVIDIA | 2021-01-08 |
| Wu Dao - Wen Yuan | Language | 6.5e+20 | 3.7e+4 | Beijing Academy of Artificial Intelligence / BAAI | 2021-01-11 |
| DLWP | Earth science | 5.7e+18 | 4.8e+2 | University of Washington,Microsoft Research | 2021-02-09 |
| MSA Transformer | Biology | 5.5e+21 | 1.5e+4 | Facebook AI Research,University of California (UC) Berkeley,New York University (NYU) | 2021-02-13 |
| Wu Dao - Wen Lan | Multimodal,Vision | 7.2e+21 | 9.7e+4 | Beijing Academy of Artificial Intelligence / BAAI | 2021-03-01 |
| Wu Dao - Wen Hui | Language,Multimodal,Video,Image generation | 1.2e+20 | 3.7e+4 | Beijing Academy of Artificial Intelligence / BAAI | 2021-03-01 |
| RFA-GATE-Gaussian-Stateful Big | Language | 7.1e+18 | 6.7e+3 | University of Washington,DeepMind,Allen Institute for AI,Hebrew University of Jerusalem,The University of Hong Kong | 2021-03-03 |
| ProteinGAN | Biology | 4.3e+18 | 4.8e+2 | Vilnius University,Chalmers University of Technology | 2021-03-04 |
| M6-T | Multimodal,Language,Vision | 5.5e+21 | 2.3e+5 | Alibaba | 2021-03-05 |
| DCTransformer (ImageNet) | Image generation | 4.4e+21 | 5.4e+4 | DeepMind | 2021-03-05 |
| AraELECTRA | Language | 2.6e+20 | 3.4e+3 | American University of Beirut | 2021-03-07 |
| Very Deep VAEs (ImageNet-64) | Image generation | 1.8e+21 | 1.8e+4 | OpenAI | 2021-03-16 |
| GLM-10B | Language | 4.9e+21 | 3.7e+4 | Tsinghua University,Beijing Academy of Artificial Intelligence / BAAI,Massachusetts Institute of Technology (MIT),Shanghai Qi Zhi institute | 2021-03-18 |
| T2R + Random Init | Language | 2.7e+19 | 4.6e+3 | University of Washington,Microsoft,DeepMind,Allen Institute for AI | 2021-03-24 |
| Transformer-C | Language | 1.9e+18 | 1.9e+3 | University of Massachusetts Amherst | 2021-04-08 |
| ProtT5-XXL-BFD | Biology | 3.7e+22 | 2.1e+5 | Technical University of Munich,Med AI Technology,NVIDIA,Oak Ridge National Laboratory,Google,Seoul National University | 2021-05-04 |
| ProtBERT-UniRef | Biology | 7.3e+21 | 2.1e+5 | Technical University of Munich,NVIDIA,Seoul National University,Google,Oak Ridge National Laboratory,Med AI Technology | 2021-05-04 |
| ProtBERT-BFD | Biology | 3.9e+22 | 4.3e+5 | Technical University of Munich,NVIDIA,Seoul National University,Google,Oak Ridge National Laboratory,Med AI Technology | 2021-05-04 |
| MedBERT | Medicine | 9.5e+18 | 4.8e+2 | Peng Cheng Laboratory,University of Texas at Houston | 2021-05-20 |
| CogView | Image generation | 2.7e+22 | 2.4e+5 | Tsinghua University,Alibaba DAMO Academy | 2021-05-26 |
| DeBERTa | Language | 2.6e+22 | 1.5e+5 | Microsoft | 2021-06-10 |
| ALIGN | Multimodal,Vision,Language | 2.6e+22 | 2.1e+5 | Google Research | 2021-06-11 |
| StyleGAN3-T | Image generation | 1.7e+21 | 4.6e+3 | NVIDIA,Aalto University | 2021-06-21 |
| StyleGAN3-R | Image generation | 2.4e+21 | 4.6e+3 | NVIDIA,Aalto University | 2021-06-21 |
| EfficientNetV2-XL | Vision | 9.6e+19 | 6.7e+3 | Google,Google Brain | 2021-06-23 |
| Fold2Seq | Biology | 1.4e+17 | 1.1e+3 | IBM,Texas A&M | 2021-06-24 |
| Adaptive Input Transformer + RD | Language | 8.6e+19 | 4.6e+3 | Microsoft Research Asia,Soochow University | 2021-06-28 |
DEQ-Transformer (Post-LN) + Jacobian Regularisation | Language | 2.9e+19 | 1.9e+3 | Carnegie Mellon University (CMU),Intel Labs | 2021-06-28 |
| GemNet-T (OC20) | Materials science | 7.3e+17 | 4.8e+2 | Technical University of Munich | 2021-07-01 |
| ERNIE 3.0 | Language | 2.2e+22 | 2.2e+5 | Baidu | 2021-07-05 |
| SEER | Vision | 1.8e+22 | 2.9e+5 | Facebook AI Research,INRIA | 2021-07-29 |
| FMMformer (2-kernel fast weight + Band20) | Language | 2.5e+16 | 3.4e+3 | University of California Los Angeles (UCLA),University of Utah | 2021-08-05 |
| YOLOX-X | Vision | 6.3e+20 | 4.6e+3 | Megvii Inc | 2021-08-06 |
| GPT-2 (1.5B, Curriculum Learning 45K) | Language | 2.4e+21 | 7.3e+4 | Microsoft | 2021-08-13 |
| ProteinLM | Biology | 1.6e+22 | 2.7e+4 | Tsinghua University,Beijing Academy of Artificial Intelligence / BAAI,Tencent | 2021-08-17 |
| $\infty$-former (SM) | Language | 1.2e+22 | 4.8e+2 | Universidade de Lisboa (ULisboa),DeepMind | 2021-09-01 |
| PermuteFormer | Language | 2.8e+18 | 4.6e+3 | Peking University | 2021-09-06 |
| NLM | Language | 2.8e+19 | 6.7e+2 | Carnegie Mellon University (CMU),University of California San Diego | 2021-09-09 |
| PLATO-XL | Language | 9.9e+21 | 1.5e+5 | Baidu | 2021-09-20 |
| Turing ULRv5 | Language | 2.9e+22 | 1.9e+5 | Microsoft | 2021-09-28 |
| AlphaFold-Multimer | Biology | 4.4e+21 | 2.7e+4 | Google DeepMind,DeepMind | 2021-10-04 |
| T0-XXL | Language | 9.2e+20 | 1.1e+5 | Hugging Face,Brown University | 2021-10-15 |
| MGK 4 heads (medium) | Language | 8.9e+18 | 1.5e+3 | FPT Software AI Center,University of California Los Angeles (UCLA),VinUniversity,Deezer Research,Rice University,University of Texas at Austin | 2021-10-16 |
| PMLM-large | Biology | 3.8e+21 | 1.4e+4 | Microsoft Research Asia,Nanyang Technological University,Xi’an Jiaotong University,Sun Yat-sen University | 2021-10-21 |
| DALL-E mini | Vision | 3.8e+19 | 1.7e+3 | Craiyon | 2021-10-26 |
| S4 | Language | 7.8e+19 | 6.1e+3 | Stanford University | 2021-10-31 |
| GPT2+CoreLM+Fine-Tuning | Language | 1.2e+20 | 4.8e+2 | Aristotle University of Thessaloniki | 2021-11-04 |
| Florence | Vision | 4.8e+22 | 3.9e+5 | Microsoft | 2021-11-22 |
| CTR-BERT | Recommendation | 6.5e+19 | 6.1e+3 | Amazon | 2021-12-06 |
| XGLM-7.5B | Language | 2.2e+22 | 1.9e+5 | Meta AI,Facebook AI Research | 2021-12-20 |
| Detic | Vision | 2.3e+19 | 1.8e+4 | Meta AI,University of Texas at Austin | 2022-01-07 |
| Primer (GPT-3 XL-like 1.9B) | Language | 2.2e+22 | 1.7e+5 | Google Brain | 2022-01-24 |
| GPT-NeoX-20B | Language | 9.3e+22 | 7.3e+4 | EleutherAI | 2022-02-09 |
| ProteinBERT | Biology | 6.5e+19 | 4.4e+2 | Hebrew University of Jerusalem,Ben-Gurion University of the Negev,Deep Trading | 2022-02-10 |
| FourCastNet | Earth science | 3.5e+20 | 4.9e+4 | NVIDIA,NERSC, Lawrence Berkeley National Laboratory,University of Michigan,Rice University,California Institute of Technology,Purdue University | 2022-02-22 |
| RQ-Transformer (3.8B params ImageNet dataset) | Vision,Image generation | 2.9e+20 | 3.0e+3 | Kakao,POSTECH | 2022-03-03 |
| GraSR | Biology | 3.8e+18 | 9.5e+2 | Shanghai Jiao Tong University,Ministry of Education of China | 2022-03-24 |
| Stable Diffusion (LDM-KL-8-G) | Image generation | 5.0e+22 | 1.9e+5 | Runway,Ludwig Maximilian University of Munich,Heidelberg University | 2022-04-13 |
| Flamingo | Multimodal,Vision,Language,Video | 2.2e+23 | 5.0e+5 | DeepMind | 2022-04-29 |
| UL2 | Language | 1.2e+23 | 1.7e+5 | Google Research,Google Brain | 2022-05-10 |
| Gato | Multimodal,Robotics,Games,Language | 4.0e+21 | 1.1e+5 | DeepMind | 2022-05-12 |
| Imagen | Image generation | 1.5e+22 | 8.3e+4 | Google Brain | 2022-05-23 |
| Tranception | Biology | 7.2e+21 | 4.9e+4 | University of Oxford,Harvard Medical School,Cohere | 2022-05-27 |
| GPT-2 Medium (FlashAttention) | Language | 8.9e+20 | 6.1e+3 | Stanford University,University at Buffalo | 2022-05-27 |
| B2T connection (16L) | Language | 2.8e+19 | 9.1e+4 | LINE Corporation,Tohoku University | 2022-06-01 |
| DITTO | Language | 3.3e+18 | 4.6e+3 | Tsinghua University,Apple,Westlake University,Chinese University of Hong Kong (CUHK) | 2022-06-06 |
| CoCa | Vision | 7.3e+22 | 6.6e+5 | Google Research | 2022-06-14 |
| YaLM | Language | 2.2e+23 | 6.1e+5 | Yandex | 2022-06-23 |
| DALL-E mega | Image generation | 2.3e+22 | 5.4e+4 | Craiyon | 2022-06-28 |
| Minerva (540B) | Language | 2.7e+24 | 3.3e+5 | 2022-06-29 | |
| BLOOM-176B | Language | 3.7e+23 | 2.9e+5 | Hugging Face,BigScience | 2022-07-11 |
| Transformer-XL + RMT | Language | 6.8e+18 | 1.5e+3 | Moscow Institute of Physics and Technology,AIRI Artificial Intelligence Research Institute | 2022-07-14 |
| ESM2-15B | Biology | 7.4e+22 | 2.9e+5 | Meta AI,New York University (NYU),Stanford University,Massachusetts Institute of Technology (MIT) | 2022-07-21 |
| ProtGPT2 | Biology | 4.1e+21 | 9.7e+4 | University of Bayreuth | 2022-07-27 |
| AlexaTM 20B | Language | 2.0e+23 | 9.7e+4 | Amazon | 2022-08-02 |
| GLM-130B | Language | 3.5e+23 | 5.8e+5 | Tsinghua University | 2022-08-04 |
| RNA-FM | Biology | 2.6e+21 | 6.1e+3 | Chinese University of Hong Kong (CUHK),Fudan University,Shanghai AI Lab,Harbin Institute of Technology,University of Electronic Science and Technology of China,Massachusetts Institute of Technology (MIT),Harvard University,Shanghai Zelixir Biotech,CUHK Shenzhen Research Institute | 2022-08-08 |
| FastSpeech 2 | Speech | 2.3e+18 | 5.7e+2 | Zhejiang University (ZJU),Microsoft Research Asia | 2022-08-08 |
| BlenderBot 3 | Language | 4.3e+23 | 9.7e+4 | McGill University,Meta AI,Mila - Quebec AI (originally Montreal Institute for Learning Algorithms) | 2022-08-10 |
| PeTriBERT | Biology | 1.0e+20 | 3.8e+3 | University of Montpellier,BionomeeX | 2022-08-13 |
| Luminous-extended | Language | 1.0e+23 | 3.9e+5 | Aleph Alpha | 2022-08-15 |
| Luminous-base | Language | 3.2e+22 | 9.7e+4 | Aleph Alpha | 2022-08-15 |
| Luminous-supreme | Language | 3.5e+23 | 3.9e+5 | Aleph Alpha | 2022-08-15 |
| Stable Diffusion 1.4 | Image generation | 5.0e+22 | 1.5e+5 | Ludwig Maximilian University of Munich | 2022-08-22 |
| PaLI | Language,Vision,Multimodal | 1.7e+23 | 3.3e+5 | 2022-09-14 | |
| AlphaTensor | Other,Games,Mathematics | 7.1e+20 | 2.7e+4 | DeepMind | 2022-10-05 |
| Decaying Fast Weights Transformer (WT-103) | Language | 7.9e+19 | 5.7e+2 | Jenni | 2022-10-09 |
| Flan-PaLM 540B | Language | 2.5e+24 | 1.7e+5 | 2022-10-20 | |
| U-PaLM (540B) | Language | 2.5e+24 | 1.7e+5 | 2022-10-20 | |
| DiffSBDD (CrossDocked) | Biology | 2.7e+20 | 7.6e+2 | Ecole Polytechnique F´ed´erale de Lausanne (EPFL),University of Cambridge,Cornell University,Chinese Academy of Mathematics and System Science,University of Rome,Microsoft Research,University of Oxford,AITHYRA Institute | 2022-10-24 |
| XY-LENTXL | Language | 7.2e+21 | 6.8e+5 | Microsoft | 2022-10-26 |
| Taiyi-Stable Diffusion | Image generation | 5.1e+22 | 2.4e+4 | IDEA CCNL | 2022-10-31 |
| Transformer-XL + PowerSGD + L-Greco | Language | 1.3e+18 | 5.3e+3 | Institute of Science and Technology Austria (ISTA),Neural Magic | 2022-10-31 |
| EVA-01 | Vision | 1.5e+22 | 9.7e+4 | Beijing Academy of Artificial Intelligence / BAAI,Huazhong University of Science and Technology,Zhejiang University (ZJU),Beijing Institute of Technology | 2022-11-14 |
| Galactica | Language,Biology | 3.2e+23 | 9.7e+4 | Meta AI | 2022-11-16 |
| CLUE | Recommendation | 3.5e+18 | 3.7e+4 | Naver Clova,Naver AI Lab | 2022-11-22 |
| Transformer + GFM | Language | 7.7e+18 | 5.3e+3 | Nanjing University | 2022-12-01 |
| ZymCTRL | Biology | 5.0e+21 | 3.7e+4 | Basecamp Research,Friedrich-Alexander-Universität,University of Girona | 2022-12-01 |
| Vega v2 | Language | 7.8e+22 | 2.4e+5 | Wuhan University,JD Explore Academy,Shanghai AI Lab,Nanyang Technological University,Washington University in St Louis,Chongqing University of Posts and Telecommunications,University of Sydney | 2022-12-04 |
| Stable Diffusion 2.1 | Image generation | 6.7e+22 | 1.9e+5 | Stability AI | 2022-12-07 |
| CaLM | Biology | 2.9e+19 | 1.2e+3 | University of Oxford | 2022-12-19 |
| OPT-IML (175B) | Language | 4.3e+23 | 9.7e+4 | Meta AI | 2022-12-22 |
| Hybrid H3-2.7B | Language | 6.5e+21 | 6.1e+3 | Stanford University,University at Buffalo | 2022-12-28 |
| SparseOPT-175B | Language | 1.6e+23 | 7.6e+2 | Institute of Science and Technology Austria (ISTA),Neural Magic | 2023-01-02 |
| DreamerV3 | Multimodal,Games | 2.2e+20 | 9.1e+3 | DeepMind,University of Toronto | 2023-01-10 |
| Nucleotide Transformer | Biology | 8.1e+21 | 9.7e+4 | NVIDIA,Technical University of Munich,InstaDeep | 2023-01-15 |
| Ankh_large | Biology | 6.5e+21 | 2.1e+4 | Technical University of Munich,Columbia University | 2023-01-16 |
| GPT-2+Active-SGD (WT2) | Language | 3.0e+17 | 5.7e+2 | University of Montreal / Université de Montréal | 2023-01-24 |
| MoLFormer-XL | Biology | 4.5e+20 | 9.1e+3 | IBM | 2023-01-25 |
| Genie-SCOPe (bio) | Biology | 1.8e+21 | 9.1e+3 | Columbia University | 2023-01-29 |
| BLIP-2 (Q-Former) | Vision,Language | 1.2e+21 | 1.2e+4 | Salesforce Research | 2023-01-30 |
| ProteinSGM | Biology | 3.0e+19 | 5.7e+2 | University of Toronto | 2023-02-04 |
| ViT-22B | Vision | 1.9e+23 | 3.3e+5 | 2023-02-10 | |
| Hyena 1.3B | Language | 4.8e+19 | 6.1e+3 | Stanford University,University of Montreal / Université de Montréal,Mila - Quebec AI (originally Montreal Institute for Learning Algorithms) | 2023-02-21 |
| Hyena-2 153M | Language | 1.9e+19 | 6.1e+3 | Stanford University,University of Montreal / Université de Montréal,Mila - Quebec AI (originally Montreal Institute for Learning Algorithms) | 2023-02-21 |
| Hyena-2 355M | Language | 3.9e+19 | 6.1e+3 | Stanford University,University of Montreal / Université de Montréal,Mila - Quebec AI (originally Montreal Institute for Learning Algorithms) | 2023-02-21 |
| Uni-Mol Molecular Model | Biology | 5.4e+18 | 3.8e+3 | Renmin University of China,DP Technology,AI for Science Institute, Beijing (AISI) | 2023-03-06 |
| Falcon-40B | Language | 2.4e+23 | 2.9e+5 | Technology Innovation Institute | 2023-03-15 |
| PanGu-Σ | Language | 4.7e+23 | 3.0e+5 | Huawei Noah's Ark Lab | 2023-03-20 |
Lightweight Fine-tuning a Pretrained Protein Language Model for Protein Secondary | Biology | 1.9e+22 | 8.6e+2 | Henan University | 2023-03-23 |
| EVA-CLIP (EVA-02-CLIP-E/14+) | Vision | 3.5e+22 | 1.6e+4 | Beijing Academy of Artificial Intelligence / BAAI,Huazhong University of Science and Technology | 2023-03-27 |
| SigLIP 400M | Vision | 4.9e+21 | 1.0e+4 | Google DeepMind | 2023-03-27 |
| SigLiT | Vision | 7.6e+19 | 1.3e+3 | Google DeepMind | 2023-03-27 |
| ERNIE-ViLG 2.0 | Image generation | 4.7e+22 | 2.4e+5 | Baidu,Wuhan University of Science and Technology | 2023-03-28 |
| VideoMAE V2 | Video | 9.7e+21 | 4.9e+4 | Nanjing University,Shenzhen Institute of Advanced Technology,Shanghai AI Lab | 2023-03-29 |
| BloombergGPT | Language | 2.4e+23 | 3.9e+5 | Bloomberg,Johns Hopkins University | 2023-03-30 |
| Pythia-12b | Language | 2.2e+22 | 1.9e+5 | EleutherAI,Booz Allen Hamilton, McLean,University of Cambridge,Indraprastha Institute of Information Technology Delhi,Stability AI,datasaur.ai,University of Amsterdam | 2023-04-03 |
| Segment Anything Model | Vision | 7.8e+21 | 1.9e+5 | Meta AI | 2023-04-05 |
| LLaVA | Multimodal,Vision,Language | 7.8e+22 | 6.1e+3 | University of Wisconsin Madison,Microsoft Research,Columbia University | 2023-04-17 |
| WizardLM-7B | Language | 4.0e+22 | 4.6e+3 | Microsoft,Peking University | 2023-04-24 |
| ruGPT-3.5 13B | Language | 1.1e+23 | 3.9e+5 | Sber | 2023-04-24 |
| Falcon-7B | Language | 6.3e+22 | 2.9e+5 | Technology Innovation Institute | 2023-04-24 |
| MosaicML Diffusion | Image generation | 1.1e+22 | 9.7e+4 | Databricks | 2023-04-28 |
| MPT-7B | Language | 4.2e+22 | 3.4e+5 | MosaicML | 2023-05-05 |
| StarCoder | Language | 8.5e+22 | 3.9e+5 | Hugging Face,ServiceNow,Northeastern University,Mila - Quebec AI (originally Montreal Institute for Learning Algorithms),Carnegie Mellon University (CMU),Johns Hopkins University,Leipzig University,ScaDS.AI,Queen Mary University of London,Roblox,Sea AI Lab,Technion - Israel Institute of Technology,Monash University,CSIRO,Data61,McGill University,Saama,University of British Columbia (UBC),Massachusetts Institute of Technology (MIT),Technical University of Munich,IBM,University of Vermont,UnfoldML,SAP,University of Notre Dame,Columbia University,New York University (NYU),University of Allahabad,Discover Dollar,Toloka,Telefonica,Stanford University,Weizmann Institute of Science,Alan Turing Institute,Wellesley College,EleutherAI,Forschungszentrum Julich | 2023-05-09 |
| InstructBLIP | Multimodal,Language,Vision | 1.9e+20 | 1.2e+4 | Salesforce Research,Hong Kong University of Science and Technology (HKUST),Nanyang Technological University | 2023-05-11 |
| ESM-GearNet | Biology | 2.1e+19 | 3.0e+3 | Mila - Quebec AI (originally Montreal Institute for Learning Algorithms),University of Montreal / Université de Montréal,IBM Research,HEC Montreal,CIFAR AI Research | 2023-05-11 |
| WeLM | Language | 2.5e+22 | 9.7e+4 | WeChat AI | 2023-05-16 |
| Baichuan1-7B | Language | 5.0e+22 | 4.9e+5 | Baichuan | 2023-06-01 |
| Polyglot-Ko-12.8B | Language | 1.3e+22 | 1.9e+5 | EleutherAI | 2023-06-04 |
| RedPajama-INCITE-7B-Base | Language | 4.1e+22 | 1.8e+6 | Together | 2023-06-06 |
| PoET | Biology | 2.3e+20 | 5.3e+3 | OpenProtein.ai | 2023-06-09 |
| MPT-30B | Language | 1.9e+23 | 6.8e+5 | MosaicML | 2023-06-22 |
| Kosmos-2 | Language,Vision,Multimodal | 4.6e+20 | 1.5e+5 | Microsoft | 2023-06-26 |
| HyenaDNA | Biology | 1.8e+21 | 6.1e+3 | Stanford University,Harvard University,Mila - Quebec AI (originally Montreal Institute for Learning Algorithms),University of Montreal / Université de Montréal | 2023-06-27 |
| Pangu-Weather | Earth science | 4.0e+22 | 1.1e+5 | Huawei | 2023-07-05 |
| xTrimoPGLM -100B | Biology | 6.2e+23 | 5.8e+5 | Tsinghua University,BioMap Research | 2023-07-06 |
| Emu1 (BAAI) | Vision,Multimodal,Language | 2.7e+21 | 9.7e+4 | Beijing Academy of Artificial Intelligence / BAAI,Tsinghua University,Peking University | 2023-07-11 |
| Llama 2-70B | Language | 8.1e+23 | 7.6e+5 | Meta AI | 2023-07-18 |
| JIANG | Language | 4.0e+22 | 7.3e+4 | K.D. Feddersen (KDF) | 2023-08-01 |
| GGNN | Biology | 7.6e+21 | 1.5e+3 | Westlake University,Tsinghua University,Toyota Technological Institute at Chicago | 2023-08-05 |
| SS-pLM | Biology | 2.3e+19 | 1.3e+3 | Nostrum Biodiscovery,Barcelona Supercomputing Center,Institucio Catalana de Recerca i Estudis Avancçats | 2023-08-06 |
| IDEFICS-80B | Multimodal,Language,Vision | 1.2e+23 | 3.9e+5 | Hugging Face | 2023-08-22 |
| ShapeMol | Biology | 2.6e+19 | 5.7e+2 | Ohio State University | 2023-08-23 |
| PeptideBERT | Biology | 4.9e+16 | 4.8e+2 | Carnegie Mellon University (CMU) | 2023-08-28 |
| Refact-1.6B | Language | 1.2e+22 | 2.6e+3 | Refact AI | 2023-08-29 |
| Swift | Robotics | 5.3e+16 | 6.7e+2 | Intel Labs | 2023-08-30 |
| TigerBot-70B | Language | 1.0e+24 | 3.9e+5 | Tigerobo | 2023-09-06 |
| DreamLLM | Multimodal,Language,Vision,Image generation | 7.5e+20 | 6.1e+4 | Xi’an Jiaotong University,Megvii Inc,Tsinghua University,Huazhong University of Science and Technology | 2023-09-20 |
| Baichuan 2-7B | Language | 1.1e+23 | 4.9e+5 | Baichuan | 2023-09-20 |
| InternLM-XComposer | Multimodal,Language,Vision | 5.3e+21 | 9.7e+4 | Shanghai AI Lab | 2023-09-26 |
| PLaMo-13B | Language | 1.2e+23 | 3.7e+5 | Preferred Networks Inc | 2023-09-28 |
| TinyLlama-1.1B (3T token checkpoint) | Language | 2.2e+22 | 1.2e+4 | Singapore University of Technology & Design | 2023-10-01 |
| Phi-1 | Language | 3.3e+20 | 6.1e+3 | Microsoft Research | 2023-10-02 |
| FoldFlow | Biology | 1.1e+20 | 3.0e+3 | McGill University,Mila - Quebec AI (originally Montreal Institute for Learning Algorithms),Dreamfold,University of Montreal / Université de Montréal,University of Oxford | 2023-10-03 |
| FinGPT-13B | Language | 1.6e+23 | 6.7e+2 | University of California Los Angeles (UCLA),Columbia University,New York University (NYU) | 2023-10-07 |
| CodeFuse-13B | Language | 3.1e+23 | 3.9e+5 | Ant Group | 2023-10-10 |
| Llemma 7B | Mathematics,Language | 1.2e+23 | 1.9e+5 | Princeton University,EleutherAI,University of Toronto,Vector Institute,University of Cambridge,Carnegie Mellon University (CMU),University of Washington | 2023-10-16 |
| Llemma 34B | Mathematics,Language | 5.4e+23 | 1.9e+5 | Princeton University,University of Toronto,Vector Institute,University of Cambridge,Carnegie Mellon University (CMU),University of Washington,EleutherAI | 2023-10-16 |
| SILC-S* (86M) | Vision | 1.0e+22 | 8.3e+4 | ETH Zurich,DeepMind,Google,Technical University of Munich | 2023-10-20 |
| CODEFUSION (Python) | Language | 7.9e+18 | 2.3e+3 | Microsoft,Microsoft Research | 2023-10-26 |
| Skywork-13B | Language | 2.5e+23 | 2.4e+5 | Kunlun Inc. | 2023-10-30 |
| Yi-34B | Language | 6.1e+23 | 9.7e+4 | 01.AI | 2023-11-02 |
| LLaVA 1.5 | Multimodal,Language,Vision | 7.8e+22 | 6.1e+3 | University of Wisconsin Madison,Microsoft Research | 2023-11-05 |
| HGRN 1B (WT 103) | Language | 2.0e+19 | 6.1e+3 | Shanghai AI Lab,Massachusetts Institute of Technology (MIT) | 2023-11-08 |
| Prithvi-100M | Earth science | 2.3e+21 | 4.9e+4 | IBM,NASA | 2023-11-08 |
| SPHINX (Llama 2 13B) | Vision,Language,Multimodal | 3.0e+22 | 2.4e+4 | Shanghai AI Lab,Chinese University of Hong Kong (CUHK),ShanghaiTech University | 2023-11-13 |
| Nemotron-3-8B | Language | 1.8e+23 | 7.8e+5 | NVIDIA | 2023-11-15 |
| Stable Video Diffusion | Image generation,Video | 6.7e+22 | 2.9e+5 | Stability AI | 2023-11-25 |
| SEA-LION V1 7B | Language | 4.3e+22 | 1.9e+5 | AI Singapore | 2023-12-01 |
| SEA-LION V1 3B | Language | 2.2e+22 | 1.8e+5 | AI Singapore | 2023-12-01 |
| Amber | Language | 4.8e+22 | 1.7e+5 | Mohamed bin Zayed University of Artificial Intelligence (MBZUAI),Petuum,University of Southern California,Carnegie Mellon University (CMU),University of Illinois Urbana-Champaign (UIUC),University of California San Diego,LLM360 | 2023-12-11 |
| VILA-13B | Multimodal | 2.3e+21 | 9.7e+4 | NVIDIA,Massachusetts Institute of Technology (MIT) | 2023-12-12 |
| Poro 34B | Language | 2.1e+23 | 4.9e+5 | High-Performance Language Technologies (HPLT),University of Turku | 2023-12-14 |
| YaYi 2.0 | Language | 4.8e+23 | 4.8e+5 | Yayi (Wenge) | 2023-12-22 |
| PLLaMa | Biology | 1.6e+23 | 6.1e+3 | University of California Santa Barbara (UCSB),University of Lincoln,Chinese Academy of Agricultural Sciences,Swedish University of Agricultural Sciences | 2024-01-03 |
Improved motif-scaffolding with SE(3) flow matching | Biology | 1.6e+19 | 1.1e+3 | University of Oxford,Massachusetts Institute of Technology (MIT),Microsoft Research AI for Science | 2024-01-08 |
| Stable Code 3B | Language | 2.1e+22 | 1.9e+5 | Stability AI | 2024-01-09 |
| InternVL | Vision,Language | 1.7e+23 | 4.9e+5 | Shanghai AI Lab,Nanjing University,The University of Hong Kong,Tsinghua University,SenseTime,University of Science and Technology of China (USTC) | 2024-01-15 |
| OmniNA | Biology | 2.5e+21 | 6.1e+3 | Tianjin Medical University | 2024-01-15 |
| StableLM-2-1.6B | Language | 1.9e+22 | 3.9e+5 | Stability AI | 2024-01-18 |
| Yi-VL-34B | Vision,Language,Multimodal | 1.9e+22 | 1.7e+5 | 01.AI | 2024-01-23 |
| ProteinStructureTransformer | Biology | 7.6e+21 | 5.3e+3 | Max Planck Institute of Biochemistry | 2024-01-26 |
| Code Llama-70B | Language | 1.3e+24 | 3.0e+5 | Meta AI | 2024-01-29 |
| LLaVA-NeXT-34B (LLaVA-1.6) | Multimodal,Language,Vision | 2.6e+20 | 2.4e+4 | University of Wisconsin Madison,ByteDance,Nanyang Technological University,University of California (UC) Berkeley | 2024-01-30 |
| CARP | Biology | 1.0e+22 | 7.3e+4 | Microsoft Research | 2024-02-06 |
| DiscDiff | Biology | 3.4e+19 | 1.5e+3 | Imperial College London | 2024-02-08 |
| PLAPT | Biology | 3.9e+22 | 3.0e+2 | Wolfram Research,ASC27,Newport High School,Sanskriti School | 2024-02-12 |
| ProtChatGPT | Biology | 7.2e+23 | 2.7e+3 | University of Technology Sydney,Zhejiang University (ZJU) | 2024-02-15 |
| Re-Dock | Biology | 7.5e+19 | 7.6e+2 | Zhejiang University (ZJU),Westlake University,University of Washington | 2024-02-21 |
| Gemma 7B | Language | 3.1e+23 | 1.1e+6 | Google DeepMind | 2024-02-21 |
| Gemma 2B | Language | 4.5e+22 | 1.4e+5 | Google DeepMind | 2024-02-21 |
| DecompDiff | Biology | 1.9e+19 | 7.6e+2 | University of Illinois Urbana-Champaign (UIUC),ByteDance,University of Chinese Academy of Sciences,Chinese Academy of Sciences,Tsinghua University | 2024-02-26 |
| ProLLaMA | Biology | 8.4e+22 | 4.6e+3 | Peking University,Peng Cheng Laboratory | 2024-02-26 |
| Nemotron-4 15B | Language | 7.5e+23 | 4.1e+6 | NVIDIA | 2024-02-27 |
| RiNALMo | Biology | 1.1e+21 | 5.3e+3 | University of Zagreb,Genome Institute of Singapore,Bioinformatics Institute | 2024-02-29 |
| MACE-MP-0 | Materials science | 8.8e+20 | 6.1e+4 | University of Cambridge,Federal Institute of Materials Research and Testing (BAM),NERSC, Lawrence Berkeley National Laboratory,University of British Columbia (UBC),Friedrich Schiller University Jena,University of Bayreuth,Fritz Haber Institute of the Max Planck Society,U. S. Naval Research Laboratory,Chemix,Daresbury Laboratory,BASF,University of South Carolina,University of Stuttgart,Uppsala University,Newcastle University,Technical University of Denmark,Aix-Marseille Université,University of Warwick,University of California Los Angeles (UCLA),InstaDeep,University of California (UC) Berkeley | 2024-03-01 |
| DeepSeek-VL-1.3B | Multimodal,Vision,Language | 8.7e+21 | 9.7e+4 | DeepSeek | 2024-03-08 |
| DeepSeek-VL-7B | Multimodal,Vision,Language | 1.0e+23 | 3.9e+5 | DeepSeek | 2024-03-08 |
| ERNIE-RNA | Biology | 2.1e+21 | 1.1e+4 | Microsoft Research,Syngentech,Tsinghua University | 2024-03-17 |
| Stable Video 3D (SV3D) | Vision,3D modeling | 5.2e+20 | 2.4e+4 | Stability AI | 2024-03-18 |
| ProstT5 | Biology | 3.1e+21 | 6.1e+3 | Technical University of Munich,Seoul National University,Institute for Advanced Study,TUM School of Life Sciences Weihenstephan | 2024-03-24 |
| TeleChat-3B | Language | 1.4e+22 | 4.9e+5 | China Telecom | 2024-04-01 |
| MobileCLIP-B (LT) | Vision | 1.3e+22 | 1.9e+5 | Apple | 2024-04-01 |
| TeleChat-7B | Language | 4.2e+22 | 4.9e+5 | China Telecom | 2024-04-01 |
| TeleChat-12B | Language | 8.6e+22 | 4.9e+5 | China Telecom | 2024-04-01 |
| Sailor-7B-Chat | Language | 1.8e+23 | 4.9e+4 | Sea AI Lab,Singapore University of Technology & Design | 2024-04-04 |
| Viking | Language | 2.6e+23 | 9.7e+5 | Silo AI,University of Turku | 2024-04-04 |
| ESM-AA | Biology | 7.3e+20 | 1.2e+4 | Peking University,Nanjing University,Tsinghua University,PharMolix | 2024-04-05 |
| YaART | Image generation | 8.2e+19 | 1.2e+5 | Yandex | 2024-04-08 |
| DDPM | Biology | 9.0e+17 | 5.7e+2 | University Paris-Saclay,Radboud University Medical Center | 2024-04-13 |
| FRED-T5-XL | Language | 2.5e+22 | 1.2e+5 | Sber | 2024-04-18 |
| SaProt | Biology | 6.2e+22 | 4.9e+4 | Zhejiang University (ZJU),Westlake University | 2024-04-19 |
| phi-3.5-Vision | Vision | 8.8e+22 | 3.4e+5 | Microsoft | 2024-04-23 |
| Phi-3.5-MoE | Language | 3.0e+23 | 6.8e+5 | Microsoft | 2024-04-23 |
| phi-3.5-mini | Language | 3.7e+22 | 6.8e+5 | Microsoft | 2024-04-23 |
| DiffPepBuilder | Biology | 7.7e+21 | 3.8e+3 | Peking University | 2024-04-30 |
| GenCast | Earth science | 8.2e+20 | 8.5e+3 | Google DeepMind | 2024-05-01 |
| OpenELM-450M | Language | 6.3e+21 | 1.7e+5 | Apple | 2024-05-02 |
| OpenELM-3B | Language | 3.4e+22 | 1.7e+5 | Apple | 2024-05-02 |
| OpenELM-270M | Language | 2.7e+21 | 9.7e+4 | Apple | 2024-05-02 |
| OpenELM-1.1B | Language | 1.1e+22 | 9.7e+4 | Apple | 2024-05-02 |
| VILA1.5-13B | Multimodal,Language,Vision | 2.3e+21 | 9.7e+4 | NVIDIA,Massachusetts Institute of Technology (MIT) | 2024-05-03 |
| xLSTM 1.4B | Language | 2.6e+18 | 7.8e+5 | Johannes Kepler University Linz | 2024-05-07 |
| AlphaFold 3 | Biology | 4.1e+22 | 1.9e+5 | Google DeepMind,Isomorphic Labs | 2024-05-08 |
| MatterSim (M3GNet - MatterSim-v1.0.0-5M) | Materials science | 1.6e+16 | 6.1e+3 | Microsoft Research AI for Science | 2024-05-10 |
| MatterSim (Grpaphomer) | Materials science | 1.1e+20 | 4.9e+4 | Microsoft Research AI for Science | 2024-05-10 |
| LBSTER | Biology | 1.1e+19 | 7.6e+2 | Prescient Design,Genentech | 2024-05-15 |
| Chameleon-34B | Multimodal,Image generation,Language,Vision | 1.6e+24 | 2.3e+6 | Facebook AI Research | 2024-05-16 |
| ProSST | Biology | 6.5e+20 | 3.8e+3 | Shanghai Jiao Tong University,Shanghai AI Lab,East China University of Science and Technology | 2024-05-17 |
| Octo-Base | Robotics | 5.8e+20 | 4.1e+4 | University of California (UC) Berkeley,Stanford University,Carnegie Mellon University (CMU),DeepMind | 2024-05-20 |
| YOLOv10-X | Vision | 1.5e+17 | 6.9e+3 | Tsinghua University | 2024-05-23 |
| Genie 2 (bio) | Biology | 3.2e+20 | 6.1e+3 | Columbia University,Rutgers University | 2024-05-24 |
| Zamba2-7B | Language | 8.8e+22 | 1.7e+5 | Zyphra | 2024-05-26 |
| Aurora | Earth science | 4.5e+21 | 2.4e+4 | Microsoft Research | 2024-05-28 |
| FoldFlow2 | Biology | 7.6e+21 | 1.5e+3 | Dreamfold,University of Montreal / Université de Montréal,McGill University,University of Oxford | 2024-05-30 |
| DRGN-AI | Biology | 6.5e+19 | 3.0e+3 | Stanford University,SLAC National Laboratory,Princeton University,Columbia University | 2024-06-02 |
| MULAN | Biology | 5.2e+20 | 1.3e+3 | AIRI Artificial Intelligence Research Institute,Skolkovo Institute of Science and Technology,Belozersky Institute of Physio-Chemical Biology,Ligand Pro | 2024-06-02 |
| Mamba2-Hybrid | Language | 1.8e+23 | 1.4e+6 | NVIDIA | 2024-06-12 |
| OpenVLA | Robotics,Vision,Language | 1.1e+23 | 4.9e+4 | Stanford University,University of California (UC) Berkeley,Toyota Research Institute,Google DeepMind,Massachusetts Institute of Technology (MIT),Physical Intelligence | 2024-06-13 |
| Ovis-7B | Multimodal,Language,Vision | 1.7e+23 | 1.7e+5 | Alibaba,Nanjing University | 2024-06-17 |
| RNA-FrameFlow | Biology | 3.7e+18 | 2.7e+3 | National University of Singapore,Prescient Design,University of Missouri,University of Cambridge | 2024-06-19 |
| Gemma 2 2B | Language | 3.1e+22 | 1.4e+5 | Google DeepMind | 2024-06-24 |
| Gemma 2 9B | Language | 4.3e+23 | 1.3e+6 | Google DeepMind | 2024-06-24 |
| JEST-L++ | Vision | 2.0e+21 | 6.8e+4 | DeepMind | 2024-06-25 |
| JEST++ | Vision | 1.9e+21 | 6.8e+4 | Google DeepMind | 2024-06-25 |
| Flexi-JEST++ | Vision | 1.3e+21 | 6.8e+4 | Google DeepMind | 2024-06-25 |
| MAP-Neo | Language | 1.9e+23 | 6.8e+5 | University of Waterloo,01.AI,Wuhan University | 2024-07-10 |
| PaliGemma | Vision | 1.1e+22 | 1.1e+5 | Google DeepMind | 2024-07-10 |
| OmniGenome | Biology | 3.4e+20 | 6.9e+3 | University of Exeter | 2024-07-15 |
| SmolLM-1.7B | Language | 1.0e+22 | 4.3e+4 | NA | 2024-07-16 |
| PepPrCLIP | Biology | 1.0e+18 | 7.6e+2 | Duke University,Cornell University,Sanford Burnham Prebys Institute | 2024-07-22 |
| AFM-server | Language | 4.3e+24 | 2.7e+6 | Apple | 2024-07-29 |
| Llama SEA-LION V2 8B | Language | 7.2e+23 | 8.5e+4 | AI Singapore | 2024-07-30 |
| EPInformer | Biology | 3.4e+17 | 7.6e+2 | The University of Hong Kong,Harvard Medical School | 2024-08-01 |
| Falcon Mamba | Language | 3.9e+23 | 3.4e+5 | Technology Innovation Institute | 2024-08-12 |
| P-LLama3 | Biology | 7.2e+23 | 2.3e+3 | University of Siena | 2024-08-12 |
| CLR_ESP | Biology | 2.1e+17 | 1.3e+2 | Kansas State University | 2024-08-16 |
| AntiFormer | Biology | 1.7e+18 | 7.6e+2 | University of Florida,Sichuan University,Shihezi University,University of Macau,University of Texas Health Science Center | 2024-08-20 |
| Pharia-1-LLM-7B | Language | 4.4e+23 | 3.4e+5 | Aleph Alpha | 2024-08-26 |
| DISTRO | Language | 7.1e+20 | 4.3e+4 | Nous Research | 2024-08-26 |
| ESMFlow | Biology | 7.6e+20 | 6.1e+3 | Massachusetts Institute of Technology (MIT) | 2024-09-02 |
| Alphaflow | Biology | 1.6e+21 | 6.1e+3 | Massachusetts Institute of Technology (MIT) | 2024-09-02 |
| OLMoE | Language | 5.2e+22 | 3.4e+5 | Allen Institute for AI,Contextual AI,University of Washington,Princeton University | 2024-09-03 |
| MPDF | Biology | 3.1e+17 | 6.7e+2 | Chinese University of Hong Kong (CUHK),Lanzhou University,Zhejiang Lab,Zhejiang University (ZJU) | 2024-09-07 |
| ALOHA Unleashed | Robotics | 3.6e+21 | 1.7e+4 | Google DeepMind | 2024-09-08 |
| MolPhenix | Biology | 4.3e+18 | 7.6e+2 | Valence Labs,University of British Columbia (UBC),Vector Institute,University of Toronto,University of Montreal / Université de Montréal,Mila - Quebec AI (originally Montreal Institute for Learning Algorithms) | 2024-09-10 |
| Novae | Biology | 1.1e+19 | 7.6e+2 | CentraleSupelec,Gustave Roussy,Université Paris Cité | 2024-09-13 |
| RNAdiffusion | Biology | 2.5e+19 | 7.6e+2 | Princeton University,Tsinghua University,Stanford University | 2024-09-15 |
| GeoSeqBuilder | Biology | 6.5e+18 | 3.1e+2 | Peking University | 2024-09-19 |
| Prithvi WxC | Earth science | 7.2e+19 | 4.9e+4 | IBM Research,University of Alabama,Stanford University,Colorado State University,Oak Ridge National Laboratory,NASA | 2024-09-20 |
| IgGM | Biology | 8.6e+20 | 6.1e+3 | Chinese Academy of Sciences,University of Chinese Academy of Sciences,Tencent | 2024-09-22 |
| TAWFN | Biology | 3.5e+18 | 6.7e+2 | Northeastern University (China) | 2024-09-23 |
| PocketGen | Biology | 2.1e+19 | 7.6e+2 | University of Science and Technology of China (USTC),Hefei Comprehensive National Science Center,Harvard University,Broad Institute,Harvard Data Science Initiative | 2024-09-23 |
| MTDP | Biology | 6.0e+14 | 2.4e+3 | Chinese University of Hong Kong (CUHK),City University of Hong Kong | 2024-09-24 |
| ProtBFN | Biology | 3.9e+22 | 8.3e+4 | InstaDeep | 2024-09-24 |
| dnaGrinder | Biology | 2.8e+19 | 1.1e+4 | Hong Kong Polytechnic University | 2024-09-24 |
| RWKV-6 (Finch) 3B | Language | 2.1e+22 | 3.7e+4 | RWKV Foundation,EleutherAI,Ohio State University,University of California Santa Barbara (UCSB),Wroclaw Tech (Wrocław University of Science and Technology),Guangdong Laboratory of Artificial Intelligence and Digital Economy (Pazhou Lab),New York University (NYU),Harvard University,Contextual AI,University of Chinese Academy of Sciences,University of California Santa Cruz,Tsinghua University,University of Edinburgh,University of British Columbia (UBC),Pennsylvania State University | 2024-09-26 |
| RWKV-5 (Eagle) 7B | Language | 5.0e+22 | 8.5e+4 | RWKV Foundation,EleutherAI,Ohio State University,University of California Santa Barbara (UCSB),Wroclaw Tech (Wrocław University of Science and Technology),Guangdong Laboratory of Artificial Intelligence and Digital Economy (Pazhou Lab),New York University (NYU),Harvard University,Contextual AI,University of Chinese Academy of Sciences,University of California Santa Cruz,Tsinghua University,University of Edinburgh,University of British Columbia (UBC),Pennsylvania State University | 2024-09-26 |
| FlexSBDD | Biology | 1.6e+19 | 7.6e+2 | University of Science and Technology of China (USTC),State Key Laboratory of Cognitive Intelligence,Princeton University | 2024-09-29 |
| scHyena | Biology | 8.6e+18 | 1.3e+3 | Korea Advanced Institute of Science and Technology (KAIST) | 2024-10-04 |
| Movie Gen Audio | Audio | 1.4e+23 | 5.1e+5 | Meta AI | 2024-10-04 |
| Movie Gen Video | Video,Vision | 1.6e+24 | 8.2e+6 | Meta AI | 2024-10-04 |
| Pyramid Flow | Video | 7.7e+21 | 9.7e+4 | Peking University,Kuaishou Technology,Beijing University of Posts and Telecommunications | 2024-10-08 |
| RDT-1B | Robotics | 4.1e+22 | 3.2e+4 | Tsinghua University | 2024-10-10 |
| CHAI-1 | Biology | 7.8e+21 | 9.7e+4 | Chai discovery | 2024-10-15 |
| Yi-Lightning | Language | 1.5e+24 | 2.7e+6 | 01.AI | 2024-10-18 |
| Allegro | Video | 6.6e+22 | 3.4e+5 | Rhymes AI | 2024-10-20 |
| Granite 3.0 2B | Language | 1.8e+23 | 1.0e+6 | IBM | 2024-10-21 |
| Granite 3.0 8B | Language | 5.8e+23 | 3.4e+5 | IBM | 2024-10-21 |
| NVLM-D 72B | Vision,Language | 3.0e+24 | 1.7e+5 | NVIDIA | 2024-10-22 |
| NVLM-H 72B | Vision,Language | 3.0e+24 | 1.7e+5 | NVIDIA | 2024-10-22 |
| NVLM-X 72B | Vision,Language | 3.0e+24 | 1.7e+5 | NVIDIA | 2024-10-22 |
| VASA-1 | Video,Audio | 4.0e+19 | 2.3e+3 | Microsoft Research Asia | 2024-10-31 |
| Uni-Med | NA | 1.4e+23 | 7.6e+2 | Tsinghua University,Beijing University of Posts and Telecommunications | 2024-11-01 |
| Fish-Speech 1.4 | Speech | 1.9e+21 | 1.1e+4 | Fish Audio | 2024-11-09 |
| NatureLM-audio | Audio | 1.4e+21 | 5.3e+3 | Earth Species Project | 2024-11-11 |
| BiRNA-BERT | Biology | 1.8e+19 | 5.3e+3 | Bangladesh University of Engineering and Technology,University of California Riverside,Carnegie Mellon University (CMU) | 2024-11-18 |
| Hymba | Language | 1.4e+22 | 9.7e+4 | NVIDIA | 2024-11-22 |
| Pleias 1.0 1.2B | Language | 3.0e+22 | 2.6e+5 | PleIAs | 2024-12-05 |
| Pleias 1.0 350m | Language | 2.7e+21 | 8.5e+4 | PleIAs | 2024-12-05 |
| NVILA 8B | Vision,Language,Multimodal,Video | 2.3e+21 | 1.7e+5 | NVIDIA,Massachusetts Institute of Technology (MIT),University of California (UC) Berkeley,University of California San Diego,University of Washington,Tsinghua University | 2024-12-05 |
| Phi-4 | Language | 9.3e+23 | 2.6e+6 | Microsoft Research | 2024-12-12 |
| F5-TTS | Speech | 4.5e+20 | 6.1e+3 | Shanghai Jiao Tong University,University of Cambridge,Geely Automobile Research Institute (Ningbo) Company | 2024-12-15 |
| Falcon3-7B | Language | 5.9e+23 | 1.4e+6 | Technology Innovation Institute | 2024-12-17 |
| SEA-LION V3 Llama3.1 70B | Language | 8.0e+24 | 2.6e+5 | AI Singapore | 2024-12-19 |
| SEA-LION V3 Llama3.1 8B | Language | 1.2e+24 | 8.5e+4 | AI Singapore | 2024-12-19 |
| SEA-LION V3 Gemma2 9B | Language | 4.5e+23 | 8.5e+4 | AI Singapore | 2024-12-19 |
| DeepSeek-V3 | Language | 3.4e+24 | 2.7e+6 | DeepSeek | 2024-12-24 |
| Cosmos-1.0- Diffusion-14B Video2World | Robotics,Vision,Video | 2.8e+24 | 1.3e+7 | NVIDIA | 2025-01-07 |
| MatterGen | Materials science | 2.7e+19 | 6.1e+3 | Microsoft Research AI for Science | 2025-01-16 |
| Eagle 2 | Vision,Robotics,Language | 4.7e+22 | 3.4e+5 | NVIDIA,Nanjing University,Tsinghua University,Hong Kong Polytechnic University,Johns Hopkins University,New York University (NYU) | 2025-01-20 |
| Zero-shot Monocular Scene Flow (ZeroMSF) | 3D modeling,Driving | 3.2e+19 | 6.1e+3 | NVIDIA,Brown University | 2025-01-20 |
| Prithvi-EO-2.0 600M | Earth science | 2.0e+22 | 1.8e+5 | IBM Research,NASA,University of Alabama,University of Iceland,Forschungszentrum Julich,Virginia Tech (Virginia Polytechnic Institute and State University),Arizona State University,Oregon State University,Boston University,University of California (UC) Berkeley,Julich Supercomputing Center | 2025-02-03 |
| Prithvi-EO-2.0 300M | Earth science | 7.1e+21 | 6.1e+4 | IBM Research,NASA,University of Alabama,University of Iceland,Forschungszentrum Julich,Virginia Tech (Virginia Polytechnic Institute and State University),Arizona State University,Oregon State University,Boston University,University of California (UC) Berkeley,Julich Supercomputing Center | 2025-02-03 |
| HAMSTER VLM | Robotics | 2.4e+21 | 6.1e+3 | NVIDIA,University of Washington,University of Southern California | 2025-02-08 |
| Brain2Qwerty | Language | 1.6e+18 | 5.7e+2 | Meta AI,Universite de Technologie de Compiègne – CNRS,Basque Center on Cognition | 2025-02-18 |
| PaliGemma 2 3B Mix 224 | Vision,Multimodal,Language | 4.0e+22 | 6.8e+4 | 2025-02-19 | |
| SigLIP 2 | Vision | 8.2e+22 | 5.5e+5 | Google DeepMind | 2025-02-20 |
| Step-Video-T2V | Video | 4.1e+24 | 6.7e+6 | StepFun | 2025-02-24 |
| Phi-4 Mini | Language | 1.0e+23 | 3.9e+5 | Microsoft | 2025-03-03 |
| Phi-4-Multimodal | Multimodal,Language,Vision,Speech | 1.2e+23 | 3.9e+5 | Microsoft | 2025-03-03 |
| Gemma 3 12B | Language,Vision,Multimodal | 8.6e+23 | 2.0e+6 | Google DeepMind | 2025-03-12 |
| Gemma 3 1B | Language | 1.2e+22 | 1.4e+5 | Google DeepMind | 2025-03-12 |
| Gemma 3 4B | Language,Vision,Multimodal | 9.6e+22 | 5.5e+5 | Google DeepMind | 2025-03-12 |
| EXAONE Deep 32B | Language | 1.3e+24 | 6.8e+5 | LG AI Research | 2025-03-16 |
| GR00T N1 2B | Robotics,Vision,Language | 7.1e+22 | 1.4e+6 | NVIDIA | 2025-03-18 |
| TxGemma 9B | Language,Biology | 4.4e+23 | 8.3e+4 | Google DeepMind,Google Research | 2025-04-08 |
| TxGemma 2B | Language,Biology | 3.2e+22 | 8.3e+4 | Google DeepMind,Google Research | 2025-04-08 |
| TxGemma 27B | Language,Biology | 2.1e+24 | 8.3e+4 | Google DeepMind,Google Research | 2025-04-08 |
| Pangu Ultra | Language | 1.1e+25 | 6.2e+6 | Huawei | 2025-04-10 |
| Nemotron-H 56B | Language | 6.7e+24 | 8.2e+6 | NVIDIA | 2025-04-14 |
| Demist-2 | Language | 4.6e+20 | 6.1e+3 | Darktrace | 2025-04-17 |
| Pleias-RAG-350m | Language | 2.7e+21 | 2.1e+4 | PleIAs | 2025-04-25 |
| Pleias-RAG-1B | Language | 3.0e+22 | 2.1e+4 | PleIAs | 2025-04-25 |
| Phi-4-Reasoning | Language | 9.3e+23 | 4.3e+4 | Microsoft | 2025-04-30 |
| Pangu Ultra MoE | Language | 3.1e+24 | 4.6e+6 | Huawei | 2025-05-07 |
| Earth-2 (cBottle-SR) | Earth science,Image generation | 1.6e+21 | 8.5e+4 | NVIDIA | 2025-05-10 |
| Falcon-H1 | Language | 3.7e+24 | 5.5e+6 | Technology Innovation Institute | 2025-05-21 |
| Pangu Pro MoE | Language | 1.3e+24 | 2.4e+6 | Huawei | 2025-05-28 |
| EXAONE 4.0 (1.2B) | Language | 8.6e+22 | 6.8e+5 | LG AI Research | 2025-07-15 |
| EXAONE 4.0 (32B) | Language | 2.7e+24 | 6.8e+5 | LG AI Research | 2025-07-15 |
| Aeneas | Vision,Multimodal,Language | 2.3e+21 | 1.7e+4 | Google DeepMind,University of Nottingham,University of Warwick,Athens University of Economics and Business,Google,University of Oxford | 2025-07-23 |
| AlphaEarth Foundations (AEF) | Earth science | 2.4e+18 | 1.7e+5 | Google DeepMind,Google | 2025-07-30 |
| Surya | Earth science | 2.9e+21 | 9.7e+4 | NASA,University of Alabama,IBM Research | 2025-08-20 |
| Teuken 7B | Language | 2.1e+23 | 3.9e+5 | OpenGPT-X,Fraunhofer Institute for Algorithms and Scientific Computing,Forschungszentrum Julich,Technische Universität Dresden | 2025-08-21 |
| Apertus 70B | Language | 6.7e+24 | 5.5e+6 | ETH Zurich,Ecole Polytechnique F´ed´erale de Lausanne (EPFL),Swiss National Supercomputing Centre (CSCS) | 2025-09-02 |
Training compute has grown even faster — around 4x/year. However, hardware efficiency (a 12x improvement in the last ten years), the adoption of lower precision formats (an 8x improvement) and longer training runs (a 4x increase) account for a roughly 2x/year decrease in power requirements relative to training compute.
Our methodology for calculating or estimating a model’s power draw during training can be found here.
Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.
Learn more about this graph
We make use of two datasets: Epoch’s AI Models dataset, which collects information on over 1000 notable AI models, as well as our Machine Learning Hardware dataset, which records information on 161 AI accelerators.
After filtering our AI models data to a subset with values for each of `Publication date`, `Training power draw (W)`, we are left with 211 AI models. Our methodology for estimating training power draw can be found here.
Analysis
We focus our analysis on frontier models released after 2010, which we define as those that were among the top 10 models by training compute at the time of their release. After filtering to frontier models released in 2010 or later, we are left with 45 observations.
We fit a log-linear model to estimate the rate of growth for frontier model power draw, and obtain a 90% confidence interval from the coefficient’s standard errors. We estimate that power draw has grown by 2.1x per year, with a 90% confidence interval of 1.9 to 2.2x.
Assumptions and limitations
In general, we define “frontier models” as those in the top 10 models ordered by training compute at the time they were published. We show the robustness of our results against other choices of top-n, as well as in the following table:
| Inclusion criteria | Number of observations | Estimate (90% confidence interval) |
|---|---|---|
| Top 5 | 29 | 2.2x per year (1.9 – 2.4) |
| Top 10 | 45 | 2.0x per year (1.7 – 2.4) |
| Top 20 | 79 | 2.1x per year (1.9 – 2.4) |
| All notable models | 175 | 1.8x per year (1.6 – 2.1) |
