The Latent
EN▾
EnglishEspañol中文PortuguêsFrançaisالعربية日本語한국어
Sign Up
NEWSBENCHMARKSDATALEARNNEWSLETTERPARTNER WITH US →
Data/Capability

Capability

MLPerf Inference Best-System Throughput

The highest valid MLPerf Inference Server throughput submitted for the Llama 2 70B 99.9% benchmark in each published round. It tracks the frontier of complete deployed systems relevant to serving large models, not per-chip efficiency, production API latency, or typical user experience.

MLPerf Inference Best-System Throughput

The best valid MLPerf Llama 2 70B Server result reached 1,016,375 tokens per second tokens per second in the 2026-03-26 round.10M1M100K10KJul '24Jan '25Jul '25Jan '26{"f":[800,420,56,16],"s":[["Best valid Llama 2 70B Server result","#2563eb"]],"p":[["2024-03-27T00:00:00.000Z","Mar '24",56,[[0,"29,526.33 tokens per second",322.95,null]],null],["2024-08-28T00:00:00.000Z","Aug '24",209.79,[[0,"82,273.2 tokens per second",268.95,null]],null],["2025-04-01T00:00:00.000Z","Apr '25",425.49,[[0,"98,443.3 tokens per second",259.49,null]],null],["2025-09-04T00:00:00.000Z","Sep '25",581.28,[[0,"153,076.32 tokens per second",236.23,null]],null],["2026-03-26T00:00:00.000Z","Mar '26",784,[[0,"1,016,375 tokens per second",136.48,null]],null]]}Best valid Llama 2 70B Server result: 1M on Mar '26
SOURCE: MLCommons MLPerf Inference results
RANGEALLYTD12M3M1M
Key takeaway

Best-system Llama 2 70B Server throughput increased from 29,526 tokens per second in MLPerf v4.0 to 1,016,375 in v6.0, reflecting rapid gains in accelerators, system scale, software, and serving optimization.

The best valid MLPerf Llama 2 70B Server result reached 1,016,375 tokens per second tokens per second in the 2026-03-26 round.

Pro API coming soon

Methodology

The chart fixes one benchmark and scenario: Llama 2 70B with the 99.9% accuracy target in the Server scenario. For each MLPerf Inference release repository from v4.0 through v6.0, the updater finds every closed- or open-division summary file for that benchmark, keeps results marked VALID, parses Completed tokens per second, and selects the highest value in the round.

Each point is dated to the corresponding MLPerf result-round publication date represented by the official repository release history. The point note records the number of candidate submissions and the winning system path, while the row source URL links directly to the winning MLPerf summary file.

The series is intentionally best-system rather than normalized per accelerator. A higher result can reflect more accelerators, a different accelerator generation, model-serving software, quantization, parallelism, and submitter tuning; it should not be interpreted as a universal throughput available from one chip or as a median production deployment.

MLPerf results are submitted and peer-reviewed under MLCommons rules, but they are benchmark submissions rather than an audit of every commercial service. Future MLPerf rounds are added by extending the reviewed round list; if Google submits a valid TPU system for this exact benchmark and scenario, it is eligible to become the best point.

Frequently asked questions

What does this chart measure?

It shows the highest valid Completed tokens per second reported for the Llama 2 70B 99.9% Server benchmark in each MLPerf Inference round. The result is a benchmarked system-throughput frontier, not a provider's average production speed.

Why is MLPerf throughput relevant to the AI industry?

Inference throughput determines how many model outputs a serving system can produce with a given hardware deployment. Tracking the best reported result shows how accelerators, interconnects, kernels, quantization, batching, and serving software are expanding the feasible scale of large-model deployment.

Does the chart compare individual GPUs or TPUs?

No. It compares the best complete systems submitted to the benchmark, and system sizes vary substantially between rounds. A large multi-accelerator result can be faster in absolute terms while being less efficient per accelerator than a smaller system.

Why use the Llama 2 70B Server benchmark?

It provides a stable large-language-model workload and a server-throughput metric across multiple MLPerf rounds. Fixing the model, accuracy target, and scenario avoids mixing benchmarks whose tokenization, quality constraints, or latency requirements would make the trend harder to interpret.

Are the results production API performance?

No. MLPerf uses a defined workload, accuracy target, load pattern, and submission configuration. Production performance also depends on prompt and output lengths, concurrency, routing, provider overhead, batching policy, service-level objectives, and the particular model implementation.

What does a Google TPU result mean if one appears?

A valid Google TPU submission would be eligible under the same best-system rule as GPU and other accelerator submissions. It would add evidence that TPU systems can compete on this standardized workload, but one vendor's benchmark result would not establish general superiority across models or operating conditions.

Related charts

  • AI Benchmark Saturation Curves
  • Largest Documented Training Compute by Release Year: China vs. U.S.
  • Frontier Model Training Compute

The Latent

AI industry news. A sister publication to The Block.

Editorial

  • Standards
  • Corrections
  • Commercial policy
  • Contact

Company

  • LEARN
  • Data
  • Benchmarks
  • Score
  • Methodology
  • About
  • Team
  • Privacy Policy
  • Terms of Service
  • Security
  • The Block
  • Add The Latent as a preferred source