The Latent Score: 2026-09-06

Release tli-2026-v5.0-2026-09-06. Methodology tli-2026-v5.0.

As of 2026-09-06, GPT-6 Astra leads The Latent Score on the point estimate at 132.9 points among 19 ranked models. Rankings describe published capability evidence, not a statistically proven winner or the best model for every workload.

This page preserves a specific publication. Its data does not change when a new release is published.

Published rankings

RankModel and evidenceScore (points)Stability range (points)Contributing results
1GPT-6 Astra132.9107.3–158.322
2Fable 5.1131.0101.4–157.520
3Opus 5125.7103.4–152.433
4Fable 5117.192.8–141.430
5Grok 4.6114.082–141.310
6Kimi K3105.878.6–13225
7GPT-5.6 Sol104.981.1–12831
8Gemini 3.7 Flash104.681.1–131.628
9Gemini 3.8 Flash104.078–131.326
10Muse Spark 1.3104.073.8–133.28
11GLM-5.394.968.9–120.623
12Qwen3.8-Max93.967.8–118.529
13GPT-5.6 Terra91.763.6–119.816
14Muse Spark 1.291.366.9–118.426
15DeepSeek V4 Pro88.862–115.523
16Sonnet 587.764–112.928
17GPT-5.6 Luna87.061.7–112.721
18DeepSeek V4 Flash84.153.2–120.65
19Qwen3.8-27B69.744.7–100.124

Methodology and limitations in this release

The scale is anchored to mean 100, SD 15 across 18 frozen reference models. It is not human IQ or a percentage. 41 benchmarks contribute across 7 domains.

One capability score combines evidence across seven equally weighted domains. Related benchmark variants share a family budget, and influence from one evaluator is capped.

Models are hard-ranked by the displayed point score. Ties use a stable model-name order. Intervals do not determine rank.

The interval reflects limited evidence, missing domains, correlated benchmark families and frozen calibration uncertainty under the scoring assumptions. It is not a promise that future scores will fall inside it 95% of the time. New conflicting evidence can widen it.

Each of seven capability domains has an equal, fixed importance budget. Missing domain abilities are conditional estimates, explicitly marked as unmeasured; individual missing benchmark results are never invented.

Selective reporting is assessed separately by withholding favorable and unfavorable results from reference models. The resulting sensitivity range has no probability interpretation and is not added to the main interval.

A model’s score and interval depend only on its own accepted evidence and the frozen methodology. Other models’ additions and updates cannot change them. A methodology revision receives a new version and can change every score.

Validation includes provider and benchmark-family holdouts, selective-reporting stress tests and replay of actual repository snapshots. Real launch history is limited and cannot establish universal future 95% coverage. This release checked 401 withheld benchmark results; normalized prediction error was 6.2% lower than a refitted legacy baseline on the same reviewed evidence. Only 2 first-recorded model arrivals across 2 providers had usable before/after evidence. Those repository dates are not verified launch dates. They do not validate launch-day coverage for every model, including Astra.

Reading this release

Compare model scores alongside their stability ranges and domain coverage. A higher point estimate does not establish superiority for every task. The linked model records disclose contributing and excluded results and their published configurations.

This release uses methodology version tli-2026-v5.0. Comparisons with another release should check both its methodology version and evidence date; a new data release does not necessarily mean the scoring method changed.

Citation and downloads

The Latent. “The Latent Score.” Release tli-2026-v5.0-2026-09-06, dataset dated 2026-09-06. Permanent release page.

Download public scores with definitions (JSON) · Data usage terms