Capability
Largest Documented Training Compute by Release Year: China vs. U.S.
Annual maximum training-compute estimate in zettaFLOP among Epoch AI notable models assigned only to China or the U.S.; 2026 is year to date. It contextualizes documented leading-run scale, but incomplete coverage and uncertain estimates preclude a national ranking.
Largest Documented Training Compute by Release Year: China vs. U.S.
- China — annual documented maximum
- U.S. — annual documented maximum
In 2026-08-24, Epoch's largest documented notable-model estimate was 20,001 zettaFLOP zettaFLOP for China and 38,700 zettaFLOP zettaFLOP for the U.S.
Methodology
Source and cohort: The chart uses Epoch AI's Notable AI Models CSV, retrieved August 27, 2026 and labeled updated August 24, 2026. Epoch recommends the notable subset for analysis. Records are limited to publication dates on or after January 1, 2023, and the last period is 2026 year-to-date rather than a completed calendar year.
Statistic: For each country and calendar release year, retain records with a nonblank Training compute (FLOP) value and select the single largest value. This is an observed annual maximum within Epoch's notable-model data, not a sum, mean, cumulative national budget, or claim about all training runs. A lower value in a later year means that year's largest documented release estimate was lower; it does not mean a country's installed compute or cumulative capability declined.
Country rule: Epoch defines Country (of organization) as the country or countries associated with the developing organization or organizations. The CSV field is split on commas, whitespace is trimmed, and duplicate identical labels are collapsed. A record qualifies only when the resulting set is exactly {China} or exactly {United States of America}. Repeated same-country labels from multiple organizations remain eligible; any record containing another country is excluded rather than split or fractionally attributed. Hong Kong remains distinct from China, matching Epoch's label, and no multinational collaboration is reassigned to either series.
Compute definition: Epoch's Training compute (FLOP) is total training compute for a model, including pretraining and fine-tuning and including pretrained base models used as components. Epoch records directly reported values where available and otherwise estimates compute from operation counts, hardware usage, cost, benchmark performance, or comparisons. For Sanity-compatible storage and readable labels, chart values divide source FLOP by 10²¹ and are displayed as zettaFLOP. The chart does not add models together or convert compute to cost, chips, energy, or capability scores.
Uncertainty: Epoch describes Confident estimates as approximately ±3×, Likely as approximately ±10×, and Speculative as approximately ±31× at 90% confidence. The selected maxima are Confident for Qwen-72B and Llama 3.1-405B; Likely for Doubao-pro and Composer 2.5; and Speculative for Gemini 1.0 Ultra, Qwen3-Max, Grok 4, and Kimi K3. Small apparent country gaps are therefore not statistically precise rankings.
Coverage: Among single-country records released from 2023 through the source cutoff, 55 of 103 China records and 73 of 206 U.S. records have nonblank training compute. Per-year denominators are recorded in the CSV notes. Missing values can hide a larger model, so each point means largest documented estimate in this dataset, not the true national frontier. Epoch's dedicated Frontier AI Models subset was not used because its two post-2023 China records both lack training-compute values; substituting them would make the comparison empty.
Presentation: A logarithmic y-axis is required because the selected estimates span roughly three orders of magnitude. Lines connect annual observations only for readability and do not interpolate training runs between release years.
Frequently asked questions
Does this chart show how much compute China and the U.S. used in total?
No. Each point is one model: the largest documented training-compute estimate among qualifying notable-model releases in that country and year. The chart never adds unrelated model training runs.
Why is this comparison relevant to the AI industry?
The maximum documented training run offers one view of the scale at which organizations in each country are developing notable models, which can inform discussion of access to advanced accelerators and large-scale infrastructure. It does not measure national investment, installed capacity, model quality or the output of the wider AI sector.
What does frontier mean here?
It is a country-year observed maximum in Epoch's notable-model data, not Epoch's global Frontier AI Models flag and not proof of the largest run that occurred. The title deliberately says largest documented training compute because missing estimates prevent a complete frontier census.
What is the source and how is each annual point selected?
The source is Epoch AI's Notable AI Models CSV. For each release year, the chart keeps models with a nonblank training-compute estimate whose organization-country set is exactly China or exactly United States of America, then selects the single highest estimate for each country. It does not sum models or fill missing values.
How are multinational labs or collaborations handled?
They are excluded from both national series whenever Epoch's country field contains more than one distinct country. Duplicate occurrences of the same country are collapsed, while Hong Kong remains a separate label from China. No model is divided between countries or assigned by headquarters guesswork.
Why can the annual maximum fall from one year to the next?
The statistic resets for each release year. A decline means the largest documented estimate among that year's qualifying releases is lower than the prior year's maximum; it does not imply that national compute capacity or model capability went backward.
Are the 2026 values comparable with full-year values?
Only as year-to-date observations. They cover releases present in Epoch's CSV through its August 24, 2026 update, and compute coverage is especially sparse in 2026: 7 of 24 qualifying China records and 4 of 29 U.S. records.
Why use a logarithmic scale?
Training-compute estimates range from 1,300 to about 500,000 zettaFLOP, where one zettaFLOP equals 10²¹ FLOP. A log scale keeps every point visible and makes multiplicative differences interpretable without implying false additive precision.