Capability
MLPerf Inference Best-System Throughput
The highest valid MLPerf Inference Server throughput submitted for the Llama 2 70B 99.9% benchmark in each published round. It tracks the frontier of complete deployed systems relevant to serving large models, not per-chip efficiency, production API latency, or typical user experience.
MLPerf Inference Best-System Throughput
The best valid MLPerf Llama 2 70B Server result reached 1,016,375 tokens per second tokens per second in the 2026-03-26 round.
Methodology
The chart fixes one benchmark and scenario: Llama 2 70B with the 99.9% accuracy target in the Server scenario. For each MLPerf Inference release repository from v4.0 through v6.0, the updater finds every closed- or open-division summary file for that benchmark, keeps results marked VALID, parses Completed tokens per second, and selects the highest value in the round.
Each point is dated to the corresponding MLPerf result-round publication date represented by the official repository release history. The point note records the number of candidate submissions and the winning system path, while the row source URL links directly to the winning MLPerf summary file.
The series is intentionally best-system rather than normalized per accelerator. A higher result can reflect more accelerators, a different accelerator generation, model-serving software, quantization, parallelism, and submitter tuning; it should not be interpreted as a universal throughput available from one chip or as a median production deployment.
MLPerf results are submitted and peer-reviewed under MLCommons rules, but they are benchmark submissions rather than an audit of every commercial service. Future MLPerf rounds are added by extending the reviewed round list; if Google submits a valid TPU system for this exact benchmark and scenario, it is eligible to become the best point.
Frequently asked questions
What does this chart measure?
It shows the highest valid Completed tokens per second reported for the Llama 2 70B 99.9% Server benchmark in each MLPerf Inference round. The result is a benchmarked system-throughput frontier, not a provider's average production speed.
Why is MLPerf throughput relevant to the AI industry?
Inference throughput determines how many model outputs a serving system can produce with a given hardware deployment. Tracking the best reported result shows how accelerators, interconnects, kernels, quantization, batching, and serving software are expanding the feasible scale of large-model deployment.
Does the chart compare individual GPUs or TPUs?
No. It compares the best complete systems submitted to the benchmark, and system sizes vary substantially between rounds. A large multi-accelerator result can be faster in absolute terms while being less efficient per accelerator than a smaller system.
Why use the Llama 2 70B Server benchmark?
It provides a stable large-language-model workload and a server-throughput metric across multiple MLPerf rounds. Fixing the model, accuracy target, and scenario avoids mixing benchmarks whose tokenization, quality constraints, or latency requirements would make the trend harder to interpret.
Are the results production API performance?
No. MLPerf uses a defined workload, accuracy target, load pattern, and submission configuration. Production performance also depends on prompt and output lengths, concurrency, routing, provider overhead, batching policy, service-level objectives, and the particular model implementation.
What does a Google TPU result mean if one appears?
A valid Google TPU submission would be eligible under the same best-system rule as GPU and other accelerator submissions. It would add evidence that TPU systems can compete on this standardized workload, but one vendor's benchmark result would not establish general superiority across models or operating conditions.