The Latent
EN▾
EnglishEspañol中文PortuguêsFrançaisالعربية日本語한국어
Sign Up
NEWSDATARESEARCHPODCASTNEWSLETTERPRESSPARTNER WITH US →
Data/Capability

Capability

Historical Token Throughput per GPU

Historical Token Throughput per GPU. Demo series compiled by The Latent research desk.

Historical Token Throughput per GPU

  • H100 FP8
  • H100 FP16
  • MI300X FP8
Historical Token Throughput per GPUDemo preview through July 2026 (tokens per second).30K20K10K0May '26Jun '26Jun '26Jun '26Jun '26H100 FP8: 25.5K on Jul '26. H100 FP16: 14.4K on Jul '26. MI300X FP8: 20.8K on Jul '26
SOURCE: The LatentUPDATED: Jul 1, 2026
ZOOMALLYTD12M3M1M
CSV
View data
DateH100 FP8H100 FP16MI300X FP8
2026-06-0125600.7314471.1720814.87
2026-07-0125466.7114413.6520758.43

API access · Pro coming soon

Demo preview through July 2026 (tokens per second).

Download CSVPro API coming soon

Methodology

Synthetic demo series generated to preview the data product layout.

Related charts

  • Historical Inference Speeds
  • Historical Token Efficiency Across Frontier Models
  • Epoch Capabilities Index (ECI)

The Latent

AI industry news. A sister publication to The Block.

Editorial

  • Standards
  • Corrections
  • Commercial policy
  • Contact

Company

  • About
  • Privacy Policy
  • Terms of Service
  • Security
  • The Block