The Latent Score: 2026-09-06

Release tli-2026-v5.1-2026-09-06. Methodology tli-2026-v5.1.

As of 2026-09-06, GPT-6 Astra leads The Latent Score on the point estimate at 132.9 points among 19 ranked models. Rankings describe published capability evidence, not a statistically proven winner or the best model for every workload.

This page preserves a specific publication. Its data does not change when a new release is published.

Published rankings

RankModel and evidenceScore (points)Stability range (points)Contributing results
1GPT-6 Astra132.9121.3–144.822
2Fable 5.1131.0123.2–138.920
3Opus 5125.7119.4–131.733
4Fable 5117.1107.8–12530
5Grok 4.6114.0102.5–126.310
6Kimi K3105.897.5–11525
7GPT-5.6 Sol104.997.8–111.931
8Gemini 3.7 Flash104.695.4–113.528
9Gemini 3.8 Flash104.092.5–115.226
10Muse Spark 1.3104.091.3–117.88
11GLM-5.394.986.7–102.723
12Qwen3.8-Max93.984.4–10429
13GPT-5.6 Terra91.781.6–10116
14Muse Spark 1.291.384.1–98.526
15DeepSeek V4 Pro88.881.5–9523
16Sonnet 587.779.5–95.928
17GPT-5.6 Luna87.077.6–95.121
18DeepSeek V4 Flash84.167.3–1005
19Qwen3.8-27B69.760–80.324

Methodology and limitations in this release

The scale is anchored to mean 100, SD 15 across 18 frozen reference models. It is not human IQ or a percentage. 41 benchmarks contribute across 7 domains.

One capability score combines evidence across seven equally weighted domains. Related benchmark variants share a family budget, and influence from one evaluator is capped.

Models are hard-ranked by the displayed point score. Ties use a stable model-name order. Intervals do not determine rank.

The bar estimates score stability under the scoring assumptions: a fixed calibration and comparable benchmark availability. It resamples benchmark families and evaluators while keeping each domain’s evidence strength fixed, then adds a frozen correction learned from missing-evidence tests. Benchmark count and domain coverage both affect that correction. It does not add a second full capability-posterior draw or uncertainty from changing the reference calibration. It is not a guarantee under selective publication, new benchmark domains or future model releases.

Each of seven capability domains has an equal, fixed importance budget. Missing domain abilities are conditional estimates, explicitly marked as unmeasured; individual missing benchmark results are never invented.

Selective reporting is assessed separately by withholding favorable and unfavorable results from reference models. The resulting sensitivity range has no probability interpretation and is not added to the main interval.

A model’s score and interval depend only on its own accepted evidence and the frozen methodology. Other models’ additions and updates cannot change them. A methodology revision receives a new version and can change every score.

Validation includes provider and benchmark-family holdouts, selective-reporting stress tests and replay of actual repository snapshots. Real launch history is limited and cannot establish universal future 95% coverage. This release checked 401 withheld benchmark results; normalized prediction error was 6.2% lower than a refitted legacy baseline on the same reviewed evidence. Only 2 first-recorded model arrivals across 2 providers had usable before/after evidence. Those repository dates are not verified launch dates. They do not validate launch-day coverage for every model, including Astra. The main score interval recovered fuller-data scores in 32 of 33 random held-out evidence cases. Selective favorable and unfavorable publication had only 18% and 12% coverage in the main bar; those risks are shown in the separate non-probabilistic reporting range. 2 historical arrivals required a conservative fallback because their earlier reference data were insufficient to train the correction. Those replays do not validate the tighter empirical bands. Calibration needs reference models from at least three providers; a narrow bar is never forced by a width cap.

Reading this release

Compare model scores alongside their stability ranges and domain coverage. A higher point estimate does not establish superiority for every task. The linked model records disclose contributing and excluded results and their published configurations.

This release uses methodology version tli-2026-v5.1. Comparisons with another release should check both its methodology version and evidence date; a new data release does not necessarily mean the scoring method changed.

Citation and downloads

The Latent. “The Latent Score.” Release tli-2026-v5.1-2026-09-06, dataset dated 2026-09-06. Permanent release page.

Download public scores with definitions (JSON) · Data usage terms