GPUs are typically faster when using tensor cores and number formats optimized for AI computing. Compared to using non-tensor FP32, TF32, tensor-FP16, and tensor-INT8 provide around 6x, 11x, and 13x greater performance on average in the aggregate performance trends. Some chips achieve even larger speedups. For example, the H100 is 59x faster in INT8 than in FP32.
| Series | Hardware | Manufacturer | Performance (OP/s) | Release date |
|---|---|---|---|---|
| FP32 (single precision) | NVIDIA Blackwell Ultra | NVIDIA | 8 x 10^13 | 2025-08-22 |
| FP32 (single precision) | NVIDIA H200 SXM | NVIDIA | 6.7 x 10^13 | 2024-11-18 |
| FP32 (single precision) | AMD Instinct MI325X | AMD | 1.6 x 10^14 | 2024-10-10 |
| FP32 (single precision) | Intel Habana Gaudi3 | Intel | 1.4 x 10^13 | 2024-09-24 |
| FP32 (single precision) | NVIDIA GB200 NVL2 (per GPU) | NVIDIA | 9 x 10^13 | 2024-06-02 |
| FP32 (single precision) | Meta MTIA v2 | Meta | 2.8 x 10^12 | 2024-04-10 |
| FP32 (single precision) | NVIDIA B200 | NVIDIA | 6.2 x 10^13 | 2024-03-18 |
| FP32 (single precision) | MTT S4000 | Moore Threads (Tencent) | 2.5 x 10^13 | 2023-12-19 |
| FP32 (single precision) | AMD Instinct MI300X | AMD | 1.6 x 10^14 | 2023-12-06 |
| FP32 (single precision) | AMD Radeon Instinct MI308X | AMD | 8.2 x 10^13 | 2023-12-06 |
| FP32 (single precision) | Sunway SW26010-Pro | Sunway | 2.8 x 10^13 | 2023-11-26 |
| FP32 (single precision) | NVIDIA HGX H20 | NVIDIA | 4.4 x 10^13 | 2023-11-09 |
| FP32 (single precision) | NVIDIA L2 PCle | NVIDIA | 2.4 x 10^13 | 2023-11-09 |
| FP32 (single precision) | NVIDIA L20 PCle | NVIDIA | 5.9 x 10^13 | 2023-09-16 |
| FP32 (single precision) | NVIDIA GH200 | NVIDIA | 6.7 x 10^13 | 2023-08-08 |
| FP32 (single precision) | MetaX MXC500 (PCIe) | MetaX | 1.5 x 10^13 | 2023-06-13 |
| FP32 (single precision) | Meta MTIA v1 | Meta | 1.6 x 10^12 | 2023-05-18 |
| FP32 (single precision) | AWS Inferentia2 | Amazon AWS | 4.8 x 10^13 | 2023-04-13 |
| FP32 (single precision) | NVIDIA H800 SXM5 | NVIDIA | 5.9 x 10^13 | 2023-03-21 |
| FP32 (single precision) | NVIDIA GH100 | NVIDIA | 6.7 x 10^13 | 2023-03-21 |
| FP32 (single precision) | NVIDIA L4 | NVIDIA | 3 x 10^13 | 2023-03-21 |
| FP32 (single precision) | NVIDIA GeForce RTX 4070 | NVIDIA | 2.9 x 10^13 | 2023-02-08 |
| FP32 (single precision) | Intel Data Center GPU Max 1100 | Intel | 2.2 x 10^13 | 2023-01-10 |
| FP32 (single precision) | Intel Data Center GPU Max 1350 | Intel | 4.4 x 10^13 | 2023-01-10 |
| FP32 (single precision) | Intel Data Center GPU Max 1550 | Intel | 5.2 x 10^13 | 2023-01-10 |
| FP32 (single precision) | AMD Radeon Instinct MI300 | AMD | 4.8 x 10^13 | 2023-01-04 |
| FP32 (single precision) | AMD Instinct MI300A | AMD | 1.2 x 10^14 | 2023-01-04 |
| FP32 (single precision) | Baidu Kunlun RG800 | Kunlunxin Baidu | 3.2 x 10^13 | 2023-01-01 |
| FP32 (single precision) | MetaX MXC500 (OAM) | MetaX | 1.8 x 10^13 | 2023-01-01 |
| FP32 (single precision) | Tianshu Zhixin Zhikai 100 | Iluvatar CoreX | 2.4 x 10^13 | 2022-12-20 |
| FP32 (single precision) | NVIDIA RTX 6000 Ada Generation | NVIDIA | 9.1 x 10^13 | 2022-12-03 |
| FP32 (single precision) | NVIDIA A800 PCIe 40 GB | NVIDIA | 1.9 x 10^13 | 2022-11-08 |
| FP32 (single precision) | AMD Radeon RX 7900 XTX | AMD | 6.1 x 10^13 | 2022-11-03 |
| FP32 (single precision) | NVIDIA A800 PCIe 80 GB | NVIDIA | 1.9 x 10^13 | 2022-11-03 |
| FP32 (single precision) | Amazon Trainium1 | Amazon AWS | 4.8 x 10^13 | 2022-10-15 |
| FP32 (single precision) | NVIDIA L40 | NVIDIA | 9.1 x 10^13 | 2022-10-13 |
| FP32 (single precision) | NVIDIA H100 PCIe | NVIDIA | 5.1 x 10^13 | 2022-09-20 |
| FP32 (single precision) | NVIDIA GeForce RTX 4090 | NVIDIA | 8.3 x 10^13 | 2022-09-20 |
| FP32 (single precision) | NVIDIA H100 SXM5 80GB | NVIDIA | 6.7 x 10^13 | 2022-09-20 |
| FP32 (single precision) | NVIDIA GeForce RTX 4080 | NVIDIA | 4.9 x 10^13 | 2022-09-20 |
| FP32 (single precision) | Biren BR100 | Biren | 2.6 x 10^14 | 2022-08-22 |
| FP32 (single precision) | NVIDIA A800 SXM | NVIDIA | 1.9 x 10^13 | 2022-08-11 |
| FP32 (single precision) | NVIDIA GeForce RTX 3090 Ti | NVIDIA | 4 x 10^13 | 2022-03-29 |
| FP32 (single precision) | Cambricon MLU370-X8 | Cambricon | 2.4 x 10^13 | 2022-03-21 |
| FP32 (single precision) | Huawei Ascend 910B | Huawei | 9.4 x 10^13 | 2022-03-17 |
| FP32 (single precision) | Baidu Kunlun R100 (75W) | Kunlunxin Baidu | 1.8 x 10^13 | 2022-01-01 |
| FP32 (single precision) | Baidu Kunlun R100 (100W) | Kunlunxin Baidu | 2 x 10^13 | 2022-01-01 |
| FP32 (single precision) | Baidu Kunlun R200 (150W) | Kunlunxin Baidu | 3.2 x 10^13 | 2022-01-01 |
| FP32 (single precision) | Baidu Kunlun R200-8F (160W) | Kunlunxin Baidu | 3.2 x 10^13 | 2022-01-01 |
| FP32 (single precision) | AMD Radeon Instinct MI250X | AMD | 4.8 x 10^13 | 2021-11-08 |
| FP32 (single precision) | Cambricon MLU370-S4/S8 | Cambricon | 1.8 x 10^13 | 2021-11-03 |
| FP32 (single precision) | Cambricon MLU370-X4 | Cambricon | 2.4 x 10^13 | 2021-11-03 |
| FP32 (single precision) | Cambricon MLU370-S8 | Cambricon | 1.8 x 10^13 | 2021-11-03 |
| FP32 (single precision) | Cambricon MLU370-S4 | Cambricon | 1.8 x 10^13 | 2021-11-03 |
| FP32 (single precision) | Tesla D1 Dojo | Tesla | 2.3 x 10^13 | 2021-08-21 |
| FP32 (single precision) | NVIDIA RTX A4000 | NVIDIA | 1.9 x 10^13 | 2021-04-12 |
| FP32 (single precision) | NVIDIA A10G | NVIDIA | 3.2 x 10^13 | 2021-04-12 |
| FP32 (single precision) | NVIDIA RTX A5000 | NVIDIA | 2.8 x 10^13 | 2021-04-12 |
| FP32 (single precision) | NVIDIA A10 PCIe | NVIDIA | 3.1 x 10^13 | 2021-04-12 |
| FP32 (single precision) | Tianshu Zhixin Tiangai 100 (BI-V100) | Iluvatar CoreX | 3.2 x 10^13 | 2021-03-31 |
| FP32 (single precision) | NVIDIA A100 SXM4 80 GB | NVIDIA | 1.9 x 10^13 | 2020-11-16 |
| FP32 (single precision) | NVIDIA RTX A6000 | NVIDIA | 3.9 x 10^13 | 2020-10-05 |
| FP32 (single precision) | NVIDIA A40 PCIe | NVIDIA | 3.7 x 10^13 | 2020-10-05 |
| FP32 (single precision) | NVIDIA GeForce RTX 3080 | NVIDIA | 3 x 10^13 | 2020-09-01 |
| FP32 (single precision) | NVIDIA GeForce RTX 3090 | NVIDIA | 3.6 x 10^13 | 2020-09-01 |
| FP32 (single precision) | NVIDIA A100 PCIe | NVIDIA | 1.9 x 10^13 | 2020-06-22 |
| FP32 (single precision) | NVIDIA A100 SXM4 40 GB | NVIDIA | 1.9 x 10^13 | 2020-05-14 |
| FP32 (single precision) | NVIDIA A100 | NVIDIA | 2 x 10^13 | 2020-03-01 |
| FP32 (single precision) | Baidu Kunlun K200 | Kunlunxin Baidu | 1.6 x 10^13 | 2020-01-01 |
| FP32 (single precision) | NVIDIA Tesla V100S PCIe 32 GB | NVIDIA | 1.6 x 10^13 | 2019-11-26 |
| FP32 (single precision) | Baidu Kunlun I | Kunlunxin Baidu | 1.6 x 10^13 | 2019-01-01 |
| FP32 (single precision) | NVIDIA Quadro RTX 4000 | NVIDIA | 7.1 x 10^12 | 2018-11-13 |
| FP32 (single precision) | NVIDIA GeForce RTX 2080 Ti 11GB | NVIDIA | 1.3 x 10^13 | 2018-09-20 |
| FP32 (single precision) | NVIDIA Tesla T4 | NVIDIA | 8.1 x 10^12 | 2018-09-13 |
| FP32 (single precision) | NVIDIA Quadro RTX 6000 | NVIDIA | 1.6 x 10^13 | 2018-08-13 |
| FP32 (single precision) | NVIDIA Quadro RTX 5000 | NVIDIA | 1.1 x 10^13 | 2018-08-13 |
| FP32 (single precision) | NVIDIA Quadro RTX 8000 | NVIDIA | 1.6 x 10^13 | 2018-08-13 |
| FP32 (single precision) | NVIDIA Tesla V100 DGXS 16 GB | NVIDIA | 1.6 x 10^13 | 2018-03-27 |
| FP32 (single precision) | NVIDIA Tesla V100 PCIe 32 GB | NVIDIA | 1.4 x 10^13 | 2018-03-27 |
| FP32 (single precision) | NVIDIA Tesla V100 DGXS 32 GB | NVIDIA | 1.6 x 10^13 | 2018-03-27 |
| FP32 (single precision) | NVIDIA Tesla V100 SXM3 32 GB | NVIDIA | 1.6 x 10^13 | 2018-03-27 |
| FP32 (single precision) | NVIDIA Tesla V100 SXM2 32 GB | NVIDIA | 1.6 x 10^13 | 2018-03-27 |
| FP32 (single precision) | NVIDIA Titan V | NVIDIA | 1.5 x 10^13 | 2017-12-01 |
| FP32 (single precision) | NVIDIA Tesla V100 SXM2 | NVIDIA | 1.6 x 10^13 | 2017-07-01 |
| FP32 (single precision) | NVIDIA Tesla V100 SXM2 16 GB | NVIDIA | 1.6 x 10^13 | 2017-06-21 |
| FP32 (single precision) | NVIDIA Tesla V100 PCIe 16 GB | NVIDIA | 1.4 x 10^13 | 2017-06-21 |
| FP32 (single precision) | NVIDIA V100 | NVIDIA | 1.6 x 10^13 | 2017-06-21 |
| FP32 (single precision) | PEZY-SC2 | PEZY | 8.2 x 10^12 | 2017-06-01 |
| FP32 (single precision) | Google TPU v2 | 3 x 10^12 | 2017-05-01 | |
| FP32 (single precision) | NVIDIA TITAN Xp | NVIDIA | 1.2 x 10^13 | 2017-04-06 |
| FP32 (single precision) | NVIDIA GeForce GTX 1080 Ti | NVIDIA | 1.1 x 10^13 | 2017-03-10 |
| FP32 (single precision) | NVIDIA Quadro P600 | NVIDIA | 1.2 x 10^12 | 2017-02-07 |
| FP32 (single precision) | NVIDIA Quadro P4000 | NVIDIA | 5.3 x 10^12 | 2017-02-06 |
| FP32 (single precision) | MT-2000 | National University of Defense Technology | 4.9 x 10^12 | 2017-01-01 |
| FP32 (single precision) | NVIDIA Quadro P5000 | NVIDIA | 8.9 x 10^12 | 2016-10-01 |
| FP32 (single precision) | NVIDIA Quadro P6000 | NVIDIA | 1.3 x 10^13 | 2016-10-01 |
| FP32 (single precision) | NVIDIA P40 | NVIDIA | 1.2 x 10^13 | 2016-09-13 |
| FP32 (single precision) | NVIDIA Tesla P4 | NVIDIA | 5.7 x 10^12 | 2016-09-13 |
| FP32 (single precision) | NVIDIA Tesla P100 PCIe 16GB | NVIDIA | 9.5 x 10^12 | 2016-06-20 |
| FP32 (single precision) | NVIDIA Tesla P100 PCIe 12GB | NVIDIA | 9.5 x 10^12 | 2016-06-20 |
| FP32 (single precision) | NVIDIA P100 PCIe 16GB | NVIDIA | 9.6 x 10^12 | 2016-06-20 |
| FP32 (single precision) | NVIDIA GeForce GTX 1080 | NVIDIA | 8.9 x 10^12 | 2016-05-27 |
| FP32 (single precision) | NVIDIA Tesla P100 SXM2 | NVIDIA | 1.1 x 10^13 | 2016-04-05 |
| FP32 (single precision) | NVIDIA P100 | NVIDIA | 9.3 x 10^12 | 2016-04-05 |
| FP32 (single precision) | NVIDIA Tesla P100 DGXS | NVIDIA | 1.1 x 10^13 | 2016-04-05 |
| FP32 (single precision) | NVIDIA M40 | NVIDIA | 6.8 x 10^12 | 2015-11-10 |
| FP32 (single precision) | NVIDIA TESLA M60 | NVIDIA | 9.6 x 10^12 | 2015-08-30 |
| FP32 (single precision) | NVIDIA Quadro M4000 | NVIDIA | 2.6 x 10^12 | 2015-06-29 |
| FP32 (single precision) | NVIDIA GeForce GTX TITAN X | NVIDIA | 6.7 x 10^12 | 2015-03-17 |
| FP32 (single precision) | NVIDIA Quadro K1200 | NVIDIA | 1.2 x 10^12 | 2015-01-28 |
| FP32 (single precision) | NVIDIA Tesla K80 | NVIDIA | 8.1 x 10^12 | 2014-11-17 |
| FP32 (single precision) | NVIDIA GeForce GTX 980 | NVIDIA | 5 x 10^12 | 2014-09-19 |
| FP32 (single precision) | NVIDIA GTX Titan Black | NVIDIA | 5.6 x 10^12 | 2014-02-18 |
| FP32 (single precision) | NVIDIA Tesla K40s | NVIDIA | 5 x 10^12 | 2013-11-22 |
| FP32 (single precision) | NVIDIA Tesla K40t | NVIDIA | 5 x 10^12 | 2013-11-22 |
| FP32 (single precision) | NVIDIA Quadro K6000 | NVIDIA | 5.2 x 10^12 | 2013-07-23 |
| FP32 (single precision) | NVIDIA GeForce GTX 780 | NVIDIA | 4.2 x 10^12 | 2013-05-23 |
| FP32 (single precision) | NVIDIA Quadro K4000 | NVIDIA | 1.2 x 10^12 | 2013-03-01 |
| FP32 (single precision) | NVIDIA GeForce GTX TITAN | NVIDIA | 4.7 x 10^12 | 2013-02-19 |
| FP32 (single precision) | NVIDIA Tesla K20c | NVIDIA | 3.5 x 10^12 | 2012-11-12 |
| FP32 (single precision) | NVIDIA Tesla K20X | NVIDIA | 3.9 x 10^12 | 2012-11-12 |
| FP32 (single precision) | NVIDIA Tesla C2050 | NVIDIA | 1 x 10^12 | 2011-07-25 |
| FP32 (single precision) | NVIDIA GeForce GTX 580 | NVIDIA | 1.6 x 10^12 | 2010-11-09 |
| FP32 (single precision) | NVIDIA GeForce GTX 280 | NVIDIA | 6.2 x 10^11 | 2008-06-16 |
| TF32 (TensorFloat-32) | NVIDIA Blackwell Ultra | NVIDIA | 1.2 x 10^15 | 2025-08-22 |
| TF32 (TensorFloat-32) | NVIDIA H200 SXM | NVIDIA | 4.9 x 10^14 | 2024-11-18 |
| TF32 (TensorFloat-32) | AMD Instinct MI325X | AMD | 6.5 x 10^14 | 2024-10-10 |
| TF32 (TensorFloat-32) | Intel Habana Gaudi3 | Intel | 4.6 x 10^14 | 2024-09-24 |
| TF32 (TensorFloat-32) | NVIDIA GB200 NVL2 (per GPU) | NVIDIA | 1.2 x 10^15 | 2024-06-02 |
| TF32 (TensorFloat-32) | NVIDIA B100 | NVIDIA | 8.8 x 10^14 | 2024-03-18 |
| TF32 (TensorFloat-32) | NVIDIA B200 | NVIDIA | 1.1 x 10^15 | 2024-03-18 |
| TF32 (TensorFloat-32) | MTT S4000 | Moore Threads (Tencent) | 5 x 10^13 | 2023-12-19 |
| TF32 (TensorFloat-32) | AMD Instinct MI300X | AMD | 6.5 x 10^14 | 2023-12-06 |
| TF32 (TensorFloat-32) | NVIDIA HGX H20 | NVIDIA | 7.4 x 10^13 | 2023-11-09 |
| TF32 (TensorFloat-32) | NVIDIA L2 PCle | NVIDIA | 4.8 x 10^13 | 2023-11-09 |
| TF32 (TensorFloat-32) | NVIDIA L20 PCle | NVIDIA | 6 x 10^13 | 2023-09-16 |
| TF32 (TensorFloat-32) | NVIDIA GH200 | NVIDIA | 4.9 x 10^14 | 2023-08-08 |
| TF32 (TensorFloat-32) | MetaX MXC500 (PCIe) | MetaX | 1.2 x 10^14 | 2023-06-13 |
| TF32 (TensorFloat-32) | AWS Inferentia2 | Amazon AWS | 1.9 x 10^14 | 2023-04-13 |
| TF32 (TensorFloat-32) | NVIDIA H800 SXM5 | NVIDIA | 4.5 x 10^14 | 2023-03-21 |
| TF32 (TensorFloat-32) | NVIDIA GH100 | NVIDIA | 4.9 x 10^14 | 2023-03-21 |
| TF32 (TensorFloat-32) | NVIDIA L4 | NVIDIA | 6 x 10^13 | 2023-03-21 |
| TF32 (TensorFloat-32) | AMD Instinct MI300A | AMD | 4.9 x 10^14 | 2023-01-04 |
| TF32 (TensorFloat-32) | MetaX MXC500 (OAM) | MetaX | 1.4 x 10^14 | 2023-01-01 |
| TF32 (TensorFloat-32) | NVIDIA A800 PCIe 40 GB | NVIDIA | 1.6 x 10^14 | 2022-11-08 |
| TF32 (TensorFloat-32) | NVIDIA A800 PCIe 80 GB | NVIDIA | 1.6 x 10^14 | 2022-11-03 |
| TF32 (TensorFloat-32) | Amazon Trainium1 | Amazon AWS | 1.9 x 10^14 | 2022-10-15 |
| TF32 (TensorFloat-32) | NVIDIA H100 PCIe | NVIDIA | 3.8 x 10^14 | 2022-09-20 |
| TF32 (TensorFloat-32) | NVIDIA GeForce RTX 4090 | NVIDIA | 8.3 x 10^13 | 2022-09-20 |
| TF32 (TensorFloat-32) | NVIDIA H100 SXM5 80GB | NVIDIA | 4.9 x 10^14 | 2022-09-20 |
| TF32 (TensorFloat-32) | Biren BR100 | Biren | 5.1 x 10^14 | 2022-08-22 |
| TF32 (TensorFloat-32) | NVIDIA A800 SXM | NVIDIA | 1.6 x 10^14 | 2022-08-11 |
| TF32 (TensorFloat-32) | NVIDIA GeForce RTX 3090 Ti | NVIDIA | 4 x 10^13 | 2022-03-29 |
| TF32 (TensorFloat-32) | AMD Radeon Instinct MI250X | AMD | 9.6 x 10^13 | 2021-11-08 |
| TF32 (TensorFloat-32) | NVIDIA A100 SXM4 80 GB | NVIDIA | 1.6 x 10^14 | 2020-11-16 |
| TF32 (TensorFloat-32) | NVIDIA A40 PCIe | NVIDIA | 7.5 x 10^13 | 2020-10-05 |
| TF32 (TensorFloat-32) | NVIDIA A100 PCIe | NVIDIA | 1.6 x 10^14 | 2020-06-22 |
| TF32 (TensorFloat-32) | NVIDIA A100 SXM4 40 GB | NVIDIA | 1.6 x 10^14 | 2020-05-14 |
| TF32 (TensorFloat-32) | NVIDIA A100 | NVIDIA | 1.6 x 10^14 | 2020-03-01 |
| TF16 (Tensor-FP16/BF16) | Huawei Ascend 920 | Huawei | 9 x 10^14 | 2025-10-01 |
| TF16 (Tensor-FP16/BF16) | NVIDIA Blackwell Ultra | NVIDIA | 2.5 x 10^15 | 2025-08-22 |
| TF16 (Tensor-FP16/BF16) | NVIDIA B300 | NVIDIA | 2.5 x 10^15 | 2025-08-15 |
| TF16 (Tensor-FP16/BF16) | Amazon Trainium2 | Amazon AWS | 6.7 x 10^14 | 2024-12-03 |
| TF16 (Tensor-FP16/BF16) | NVIDIA H200 SXM | NVIDIA | 9.9 x 10^14 | 2024-11-18 |
| TF16 (Tensor-FP16/BF16) | Intel Habana Gaudi3 | Intel | 1.7 x 10^15 | 2024-09-24 |
| TF16 (Tensor-FP16/BF16) | Maia 100 (M100) | Microsoft | 8 x 10^14 | 2024-08-27 |
| TF16 (Tensor-FP16/BF16) | NVIDIA GB200 NVL2 (per GPU) | NVIDIA | 2.5 x 10^15 | 2024-06-02 |
| TF16 (Tensor-FP16/BF16) | Google TPU v6e Trillium | 9.2 x 10^14 | 2024-05-14 | |
| TF16 (Tensor-FP16/BF16) | Meta MTIA v2 | Meta | 1.8 x 10^14 | 2024-04-10 |
| TF16 (Tensor-FP16/BF16) | NVIDIA B100 | NVIDIA | 1.8 x 10^15 | 2024-03-18 |
| TF16 (Tensor-FP16/BF16) | NVIDIA B200 | NVIDIA | 2.2 x 10^15 | 2024-03-18 |
| TF16 (Tensor-FP16/BF16) | MTT S4000 | Moore Threads (Tencent) | 1 x 10^14 | 2023-12-19 |
| TF16 (Tensor-FP16/BF16) | Google TPU v5p | 4.6 x 10^14 | 2023-12-06 | |
| TF16 (Tensor-FP16/BF16) | AMD Instinct MI300X | AMD | 1.3 x 10^15 | 2023-12-06 |
| TF16 (Tensor-FP16/BF16) | NVIDIA HGX H20 | NVIDIA | 1.5 x 10^14 | 2023-11-09 |
| TF16 (Tensor-FP16/BF16) | NVIDIA L2 PCle | NVIDIA | 9.6 x 10^13 | 2023-11-09 |
| TF16 (Tensor-FP16/BF16) | NVIDIA L20 PCle | NVIDIA | 1.2 x 10^14 | 2023-09-16 |
| TF16 (Tensor-FP16/BF16) | Google TPU v5e | 2 x 10^14 | 2023-08-29 | |
| TF16 (Tensor-FP16/BF16) | NVIDIA GH200 | NVIDIA | 9.9 x 10^14 | 2023-08-08 |
| TF16 (Tensor-FP16/BF16) | Meta MTIA v1 | Meta | 5.1 x 10^13 | 2023-05-18 |
| TF16 (Tensor-FP16/BF16) | AWS Inferentia2 | Amazon AWS | 1.9 x 10^14 | 2023-04-13 |
| TF16 (Tensor-FP16/BF16) | NVIDIA H800 SXM5 | NVIDIA | 9.9 x 10^14 | 2023-03-21 |
| TF16 (Tensor-FP16/BF16) | NVIDIA GH100 | NVIDIA | 9.9 x 10^14 | 2023-03-21 |
| TF16 (Tensor-FP16/BF16) | NVIDIA L4 | NVIDIA | 1.2 x 10^14 | 2023-03-21 |
| TF16 (Tensor-FP16/BF16) | AMD Instinct MI300A | AMD | 9.8 x 10^14 | 2023-01-04 |
| TF16 (Tensor-FP16/BF16) | NVIDIA A800 PCIe 40 GB | NVIDIA | 3.1 x 10^14 | 2022-11-08 |
| TF16 (Tensor-FP16/BF16) | NVIDIA A800 PCIe 80 GB | NVIDIA | 3.1 x 10^14 | 2022-11-03 |
| TF16 (Tensor-FP16/BF16) | Amazon Trainium1 | Amazon AWS | 1.9 x 10^14 | 2022-10-15 |
| TF16 (Tensor-FP16/BF16) | NVIDIA L40 | NVIDIA | 1.8 x 10^14 | 2022-10-13 |
| TF16 (Tensor-FP16/BF16) | NVIDIA H100 PCIe | NVIDIA | 7.6 x 10^14 | 2022-09-20 |
| TF16 (Tensor-FP16/BF16) | NVIDIA GeForce RTX 4090 | NVIDIA | 3.3 x 10^14 | 2022-09-20 |
| TF16 (Tensor-FP16/BF16) | NVIDIA H100 SXM5 80GB | NVIDIA | 9.9 x 10^14 | 2022-09-20 |
| TF16 (Tensor-FP16/BF16) | Biren BR100 | Biren | 1 x 10^15 | 2022-08-22 |
| TF16 (Tensor-FP16/BF16) | NVIDIA A800 SXM | NVIDIA | 3.1 x 10^14 | 2022-08-11 |
| TF16 (Tensor-FP16/BF16) | Intel Habana Gaudi2 | Intel | 4.5 x 10^14 | 2022-05-10 |
| TF16 (Tensor-FP16/BF16) | NVIDIA GeForce RTX 3090 Ti | NVIDIA | 1.6 x 10^14 | 2022-03-29 |
| TF16 (Tensor-FP16/BF16) | AMD Radeon Instinct MI250X | AMD | 3.8 x 10^14 | 2021-11-08 |
| TF16 (Tensor-FP16/BF16) | Tesla D1 Dojo | Tesla | 3.6 x 10^14 | 2021-08-21 |
| TF16 (Tensor-FP16/BF16) | Google TPU v4 | 2.8 x 10^14 | 2021-05-20 | |
| TF16 (Tensor-FP16/BF16) | NVIDIA A100 SXM4 80 GB | NVIDIA | 3.1 x 10^14 | 2020-11-16 |
| TF16 (Tensor-FP16/BF16) | NVIDIA A40 PCIe | NVIDIA | 1.5 x 10^14 | 2020-10-05 |
| TF16 (Tensor-FP16/BF16) | NVIDIA A100 PCIe | NVIDIA | 3.1 x 10^14 | 2020-06-22 |
| TF16 (Tensor-FP16/BF16) | NVIDIA A100 SXM4 40 GB | NVIDIA | 3.1 x 10^14 | 2020-05-14 |
| TF16 (Tensor-FP16/BF16) | NVIDIA A100 | NVIDIA | 3.1 x 10^14 | 2020-03-01 |
| TF16 (Tensor-FP16/BF16) | Google TPU v4i | 1.4 x 10^14 | 2020-01-01 | |
| TF16 (Tensor-FP16/BF16) | NVIDIA Tesla V100S PCIe 32 GB | NVIDIA | 1.3 x 10^14 | 2019-11-26 |
| TF16 (Tensor-FP16/BF16) | NVIDIA GeForce RTX 2080 Ti 11GB | NVIDIA | 1.1 x 10^14 | 2018-09-20 |
| TF16 (Tensor-FP16/BF16) | Google TPU v3 | 1.2 x 10^14 | 2018-05-18 | |
| TF16 (Tensor-FP16/BF16) | NVIDIA Tesla V100 DGXS 16 GB | NVIDIA | 1.2 x 10^14 | 2018-03-27 |
| TF16 (Tensor-FP16/BF16) | NVIDIA Tesla V100 PCIe 32 GB | NVIDIA | 1.1 x 10^14 | 2018-03-27 |
| TF16 (Tensor-FP16/BF16) | NVIDIA Tesla V100 DGXS 32 GB | NVIDIA | 1.2 x 10^14 | 2018-03-27 |
| TF16 (Tensor-FP16/BF16) | NVIDIA Tesla V100 SXM3 32 GB | NVIDIA | 1.2 x 10^14 | 2018-03-27 |
| TF16 (Tensor-FP16/BF16) | NVIDIA Tesla V100 SXM2 32 GB | NVIDIA | 1.2 x 10^14 | 2018-03-27 |
| TF16 (Tensor-FP16/BF16) | NVIDIA Tesla V100 SXM2 | NVIDIA | 1.2 x 10^14 | 2017-07-01 |
| TF16 (Tensor-FP16/BF16) | NVIDIA Tesla V100 SXM2 16 GB | NVIDIA | 1.2 x 10^14 | 2017-06-21 |
| TF16 (Tensor-FP16/BF16) | NVIDIA Tesla V100 PCIe 16 GB | NVIDIA | 1.1 x 10^14 | 2017-06-21 |
| TF16 (Tensor-FP16/BF16) | NVIDIA V100 | NVIDIA | 1.2 x 10^14 | 2017-06-21 |
| TF16 (Tensor-FP16/BF16) | Google TPU v2 | 4.5 x 10^13 | 2017-05-01 | |
| INT8 | NVIDIA Blackwell Ultra | NVIDIA | 1.6 x 10^14 | 2025-08-22 |
| INT8 | NVIDIA H200 SXM | NVIDIA | 2 x 10^15 | 2024-11-18 |
| INT8 | AMD Instinct MI325X | AMD | 2.6 x 10^15 | 2024-10-10 |
| INT8 | NVIDIA GB200 NVL2 (per GPU) | NVIDIA | 5 x 10^15 | 2024-06-02 |
| INT8 | Google TPU v6e Trillium | 1.8 x 10^15 | 2024-05-14 | |
| INT8 | Meta MTIA v2 | Meta | 3.5 x 10^14 | 2024-04-10 |
| INT8 | NVIDIA B100 | NVIDIA | 3.5 x 10^15 | 2024-03-18 |
| INT8 | NVIDIA B200 | NVIDIA | 4.5 x 10^15 | 2024-03-18 |
| INT8 | MTT S4000 | Moore Threads (Tencent) | 2 x 10^14 | 2023-12-19 |
| INT8 | Google TPU v5p | 9.2 x 10^14 | 2023-12-06 | |
| INT8 | AMD Instinct MI300X | AMD | 2.6 x 10^15 | 2023-12-06 |
| INT8 | NVIDIA HGX H20 | NVIDIA | 3 x 10^14 | 2023-11-09 |
| INT8 | NVIDIA L2 PCle | NVIDIA | 1.9 x 10^14 | 2023-11-09 |
| INT8 | NVIDIA L20 PCle | NVIDIA | 2.4 x 10^14 | 2023-09-16 |
| INT8 | Google TPU v5e | 3.9 x 10^14 | 2023-08-29 | |
| INT8 | NVIDIA GH200 | NVIDIA | 2 x 10^15 | 2023-08-08 |
| INT8 | MetaX MXC500 (PCIe) | MetaX | 4.8 x 10^14 | 2023-06-13 |
| INT8 | MetaX MXN100 | MetaX | 1.6 x 10^14 | 2023-06-07 |
| INT8 | Meta MTIA v1 | Meta | 1 x 10^14 | 2023-05-18 |
| INT8 | AWS Inferentia2 | Amazon AWS | 3.8 x 10^14 | 2023-04-13 |
| INT8 | NVIDIA H800 SXM5 | NVIDIA | 2 x 10^15 | 2023-03-21 |
| INT8 | NVIDIA GH100 | NVIDIA | 2 x 10^15 | 2023-03-21 |
| INT8 | NVIDIA L4 | NVIDIA | 2.4 x 10^14 | 2023-03-21 |
| INT8 | NVIDIA GeForce RTX 4070 | NVIDIA | 2.3 x 10^14 | 2023-02-08 |
| INT8 | AMD Instinct MI300A | AMD | 2 x 10^15 | 2023-01-04 |
| INT8 | Baidu Kunlun RG800 | Kunlunxin Baidu | 2.6 x 10^14 | 2023-01-01 |
| INT8 | MetaX MXC500 (OAM) | MetaX | 5.6 x 10^14 | 2023-01-01 |
| INT8 | Tianshu Zhixin Zhikai 100 | Iluvatar CoreX | 3.8 x 10^14 | 2022-12-20 |
| INT8 | Amazon Trainium1 | Amazon AWS | 4.2 x 10^14 | 2022-10-15 |
| INT8 | GroqChip LPU v1 | NA | 7.5 x 10^14 | 2022-10-15 |
| INT8 | NVIDIA L40 | NVIDIA | 3.6 x 10^14 | 2022-10-13 |
| INT8 | NVIDIA GeForce RTX 4090 | NVIDIA | 6.6 x 10^14 | 2022-09-20 |
| INT8 | NVIDIA H100 SXM5 80GB | NVIDIA | 2 x 10^15 | 2022-09-20 |
| INT8 | NVIDIA GeForce RTX 4080 | NVIDIA | 3.9 x 10^14 | 2022-09-20 |
| INT8 | Biren BR100 | Biren | 2 x 10^15 | 2022-08-22 |
| INT8 | NVIDIA GeForce RTX 3090 Ti | NVIDIA | 3.2 x 10^14 | 2022-03-29 |
| INT8 | Cambricon MLU370-X8 | Cambricon | 2.6 x 10^14 | 2022-03-21 |
| INT8 | Baidu Kunlun R100 (75W) | Kunlunxin Baidu | 1.3 x 10^14 | 2022-01-01 |
| INT8 | Baidu Kunlun R100 (100W) | Kunlunxin Baidu | 1.7 x 10^14 | 2022-01-01 |
| INT8 | Baidu Kunlun R200 (150W) | Kunlunxin Baidu | 2.6 x 10^14 | 2022-01-01 |
| INT8 | Baidu Kunlun R200-8F (160W) | Kunlunxin Baidu | 2.6 x 10^14 | 2022-01-01 |
| INT8 | AMD Radeon Instinct MI250X | AMD | 3.8 x 10^14 | 2021-11-08 |
| INT8 | Cambricon MLU370-S4/S8 | Cambricon | 1.9 x 10^14 | 2021-11-03 |
| INT8 | Cambricon MLU370-X4 | Cambricon | 2.6 x 10^14 | 2021-11-03 |
| INT8 | Cambricon MLU370-S8 | Cambricon | 1.9 x 10^14 | 2021-11-03 |
| INT8 | Cambricon MLU370-S4 | Cambricon | 1.9 x 10^14 | 2021-11-03 |
| INT8 | Baidu Kunlun II | Kunlunxin Baidu | 2.6 x 10^14 | 2021-08-18 |
| INT8 | Google TPU v4 | 2.8 x 10^14 | 2021-05-20 | |
| INT8 | NVIDIA A10 PCIe | NVIDIA | 2.5 x 10^14 | 2021-04-12 |
| INT8 | Tianshu Zhixin Tiangai 100 (BI-V100) | Iluvatar CoreX | 2.6 x 10^14 | 2021-03-31 |
| INT8 | Cambricon MLU290-M5 | Cambricon | 5.1 x 10^14 | 2021-01-21 |
| INT8 | NVIDIA A100 SXM4 80 GB | NVIDIA | 6.2 x 10^14 | 2020-11-16 |
| INT8 | NVIDIA A40 PCIe | NVIDIA | 3 x 10^14 | 2020-10-05 |
| INT8 | NVIDIA GeForce RTX 3090 | NVIDIA | 2.8 x 10^14 | 2020-09-01 |
| INT8 | NVIDIA A100 PCIe | NVIDIA | 6.2 x 10^14 | 2020-06-22 |
| INT8 | NVIDIA A100 SXM4 40 GB | NVIDIA | 6.2 x 10^14 | 2020-05-14 |
| INT8 | NVIDIA A100 | NVIDIA | 6.2 x 10^14 | 2020-03-01 |
| INT8 | Baidu Kunlun K200 | Kunlunxin Baidu | 2.6 x 10^14 | 2020-01-01 |
| INT8 | T-Head Hanguang 800 | Alibaba | 8.2 x 10^14 | 2019-09-25 |
| INT8 | Huawei Ascend 910 | Huawei | 5.1 x 10^14 | 2019-08-23 |
| INT8 | Huawei Ascend 910 (320 TFLOPS) | Huawei | 6.4 x 10^14 | 2019-08-23 |
| INT8 | Cambricon MLU270-F4 | Cambricon | 1.3 x 10^14 | 2019-06-20 |
| INT8 | Cambricon MLU270-S4 | Cambricon | 1.3 x 10^14 | 2019-06-20 |
| INT8 | Baidu Kunlun I | Kunlunxin Baidu | 2.6 x 10^14 | 2019-01-01 |
| INT8 | Huawei Ascend 310 | Huawei | 1.6 x 10^13 | 2018-10-10 |
| INT8 | NVIDIA GeForce RTX 2080 Ti 11GB | NVIDIA | 2.3 x 10^14 | 2018-09-20 |
| INT8 | NVIDIA Tesla P4 | NVIDIA | 2.2 x 10^13 | 2016-09-13 |
| INT8 | Google TPU v1 | 9.2 x 10^13 | 2015-05-15 |
These improvements account for about half of the overall performance trend improvement since their introduction. Models trained with lower precision formats have become common, especially tensor-FP16, as developers take advantage of this boost in performance.
Epoch's work is free to use, distribute, and reproduce provided the source and authors are credited under the Creative Commons BY license.
Learn more about this graph
We estimate the individual per-format trends, then use these to estimate the average performance improvement compared to the non-tensor FP32 trend. This average improvement is evaluated across the limits of the shorter data series, so typically begins around 2016. Note that individual chips can vary significantly in how much they benefit from different number formats.
Analysis
We perform separate log-linear regressions, for each numeric format, between computational performance and release date. The results from these regressions are used to estimate the average performance improvement relative to FP32. Confidence intervals for this difference are derived from the standard errors of the regression averages.
Assumptions
We assume that the aggregate performance trends across ML accelerators in our dataset are representative of hardware used for ML. This is an imperfect assumption, because a few leading hardware models are used more often. We believe that including other ML hardware is nevertheless useful to better estimate performance trends, but if flagship ML hardware follows different trends to ML hardware more broadly, these results will be less accurate.
For average performance improvement to reflect performance improvement at a given time, we must assume that the individual performance trends are close to parallel. So far, this appears to have held true.

