The Latent
EN▾
EnglishEspañol中文PortuguêsFrançaisالعربية日本語한국어
Sign Up
NEWSDATARESEARCHPODCASTNEWSLETTERPRESSPARTNER WITH US →
Data/Capability

Capability

Historical Token Throughput per GPU

Historical Token Throughput per GPU. Demo series compiled by The Latent research desk.

Historical Token Throughput per GPU

  • H100 FP8
  • H100 FP16
  • MI300X FP8
Historical Token Throughput per GPUDemo preview through July 2026 (tokens per second).30K20K10K0Aug '25Oct '25Dec '25Feb '26Apr '26Jun '26H100 FP8: 25.5K on Jul '26. H100 FP16: 14.4K on Jul '26. MI300X FP8: 20.8K on Jul '26
SOURCE: The LatentUPDATED: Jul 1, 2026
ZOOMALLYTD12M3M1M
CSV
View data
DateH100 FP8H100 FP16MI300X FP8
2025-08-0120069.3811352.6316341.1
2025-09-0120118.7811370.5716352.71
2025-10-0122483.212706.0218271.7
2025-11-0122688.2812836.4218480.2
2025-12-0122293.2412607.1418141.94
2026-01-0122815.4712908.6718584.56
2026-02-0121223.9212010.5617295.3
2026-03-0121860.1412353.2417763.69
2026-04-0123606.213347.1719203.15
2026-05-0123700.5513399.2419276.44
2026-06-0125600.7314471.1720814.87
2026-07-0125466.7114413.6520758.43

API access · Pro coming soon

Demo preview through July 2026 (tokens per second).

Download CSVPro API coming soon

Methodology

Synthetic demo series generated to preview the data product layout.

Related charts

  • Historical Inference Speeds
  • Historical Token Efficiency Across Frontier Models
  • Epoch Capabilities Index (ECI)

The Latent

AI industry news. A sister publication to The Block.

Editorial

  • Standards
  • Corrections
  • Commercial policy
  • Contact

Company

  • About
  • Privacy Policy
  • Terms of Service
  • Security
  • The Block