The Latent
EN▾
EnglishEspañol中文PortuguêsFrançaisالعربية日本語한국어
Sign Up
NEWSBENCHMARKSDATALEARNNEWSLETTERPARTNER WITH US →
Data/Capability

Capability

AI Benchmark Saturation Curves

Running record accuracy on four version-filtered Epoch AI benchmarks, indexed by model release date. The curves show when evaluations lose headroom for distinguishing frontier systems—important context for benchmark replacement—without treating percentage points across tasks as equivalent capability.

AI Benchmark Saturation Curves

As of the Epoch AI export updated 2026-08-13, the latest observed records were 94.82% on GPQA Diamond, 98.13% on MATH Level 5, 100% on OTIS Mock AIME 2024–2025, and 83.47% on Epoch's SWE-bench Verified v2.100%75%50%25%0%{"f":[800,420,56,16],"s":[["GPQA Diamond","#6D5DFC"]],"p":[["2026-08-13T00:00:00.000Z","Aug '26",56,[[0,"94.82%",34.84,null]],null]]}GPQA Diamond: 94.8% on Aug '26
SOURCE: Epoch AI Benchmarking Hub
RANGEALLYTD12M3M1M
Key takeaway

The compatible Epoch observations show clear ceiling effects: MATH Level 5 reached 98.13%, OTIS Mock AIME reached 100%, and GPQA Diamond reached 94.82%. Those levels leave little score headroom for separating future models on these setups. Epoch's SWE-bench Verified v2 record was 83.47%, but its compatible history is much shorter.

As of the Epoch AI export updated 2026-08-13, the latest observed records were 94.82% on GPQA Diamond, 98.13% on MATH Level 5, 100% on OTIS Mock AIME 2024–2025, and 83.47% on Epoch's SWE-bench Verified v2.

Pro API coming soon

Methodology

The sole observation source is Epoch AI's downloadable LLM Benchmark Data ZIP, retrieved from https://epoch.ai/data/benchmark_data.zip on 2026-08-27. Epoch states that its data may be used, distributed, and reproduced with attribution under Creative Commons Attribution 4.0. This derivative contains model-level scores and identifiers only, not benchmark questions or answers.

The y-axis is percent correct. Each included source file reports a mean_score bounded from 0 to 1 and representing accuracy or resolved-task rate, so the only transformation is multiplication by 100. No min-max normalization, random-baseline correction, difficulty adjustment, cross-benchmark averaging, or composite score is used. A percentage point remains benchmark-specific even though the four lines share a display axis.

The x-axis is the model Release date supplied by Epoch, not the evaluation date. Historical models can therefore appear before Epoch ran the evaluation. Eligible release dates are from 2023-01-01 through 2026-08-27 inclusive. Rows without a release date, mean score, or evaluation start time are excluded.

For each benchmark and release date, the highest eligible mean_score is retained. Dates are then sorted and a point is emitted only when its score strictly exceeds every earlier eligible score in that same benchmark. The intended rendering is a right-continuous step-after line. The CSV does not add synthetic monthly observations or extend a record to the 2026-08-27 source cutoff.

Compatibility is enforced separately by series. GPQA Diamond uses evaluations started on or after its 2026-02-20 v1.0.6 answer-format update. OTIS Mock AIME uses evaluations started on or after 2025-09-01, when Epoch says it replaced its answer-aware model grader with answer extraction plus exact matching. SWE-bench Verified uses evaluations started on or after Epoch's 2026-02-12 v2.0.0 overhaul. These filters sacrifice historical coverage rather than splice materially different procedures.

MATH Level 5 uses Epoch's dashboard mean score under the model_graded_equiv scorer described on the benchmark page. Epoch evaluates 1,324 level-5 test questions and warns that overlap with mathematical fine-tuning data may inflate scores. Its curve is retained because saturation is the subject of the chart, but it must not be read as contamination-adjusted capability.

Uncertainty is not used to decide records: the running frontier follows Epoch's mean point estimate. The source standard error in percentage points is preserved in every CSV note. Very small record increments can therefore be statistically indistinguishable; for example, OTIS Mock AIME's 97.78% and 97.80% points are both shown because the latter is a strict point-estimate record.

FrontierMath was excluded because the 2026-08-27 archive's frontiermath.csv does not expose a benchmark-version field while Epoch documents a major v2 correction affecting 42% of problems on 2026-06-12. External leaderboards were also excluded because their prompts, scaffolds, scorers, sample sets, and licensing vary. Neither incompatible versions nor heterogeneous external metrics are silently joined.

Frequently asked questions

What does a benchmark saturation curve show?

It shows the highest score observed among eligible models released up to each record date. A line nearing 100% indicates that the benchmark has little remaining headroom, not that the model has achieved 100% of general intelligence or capability.

Why does benchmark saturation matter to the AI industry?

Benchmarks help labs, customers and researchers compare model releases. When leading scores bunch near the ceiling, a test has less room to distinguish further progress, making harder, cleaner or differently scoped evaluations more important. Saturation does not show that the underlying domain is solved.

Are scores from different benchmarks directly comparable?

Only in the limited sense that all four are percentages of items or tasks passed. A 90% score on GPQA is not equivalent to 90% on SWE-bench or MATH. The shared axis is for seeing each benchmark's own approach to its ceiling, and no cross-benchmark aggregate is calculated.

Why are these step lines instead of every model result?

Each point is a strict new same-benchmark record by model release date. A step-after line makes the observed frontier and shrinking ceiling headroom legible without lower-scoring models obscuring it.

What is the source and how are record points selected?

Scores come from Epoch AI's downloadable LLM Benchmark Data archive. Within each version-compatible series and model release date, the chart retains the highest eligible mean score, converts the 0–1 value to percent, and plots a point only when it strictly exceeds all earlier eligible scores in that benchmark.

Why does GPQA Diamond start in late 2024?

Epoch changed the parsed answer format in benchmark v1.0.6 on February 20, 2026. To avoid mixing prompt/parser versions, this chart uses only runs started on or after that date. The earliest compatible record in the export belongs to a model released on December 17, 2024.

Why does SWE-bench Verified start in 2025?

Epoch introduced a major v2.0.0 methodology update on February 12, 2026 and says its default graph only displays v2 results. This chart applies the same cutoff to evaluation start times; the earliest qualifying record is a model released on April 16, 2025.

Does MATH Level 5's near-100% record prove the benchmark is cleanly solved?

No. It demonstrates observed saturation under Epoch's evaluation setup. Epoch explicitly cautions that models may have encountered overlapping mathematical content during training, so contamination may contribute to the result.

Why is FrontierMath omitted?

Epoch released a major corrected v2 in June 2026, but the downloadable CSV used here lacks an explicit version column. Without a reliable row-level version filter, combining its historical scores would risk splicing incompatible benchmark forms.

Can this data be redistributed?

Yes, for these Epoch-provided observations: Epoch's archive README and licensing page state Creative Commons Attribution 4.0 terms. Attribution to Epoch AI is required. Benchmark questions and answers remain the property of their creators and are not included here.

Related charts

  • Largest Documented Training Compute by Release Year: China vs. U.S.
  • Frontier Model Training Compute
  • Maximum API Model Context Window Over Time

The Latent

AI industry news. A sister publication to The Block.

Editorial

  • Standards
  • Corrections
  • Commercial policy
  • Contact

Company

  • LEARN
  • Data
  • Benchmarks
  • Score
  • Methodology
  • About
  • Team
  • Privacy Policy
  • Terms of Service
  • Security
  • The Block
  • Add The Latent as a preferred source