The Latent Score: 2026-09-06
Release tli-2026-v5.0-2026-09-06. Methodology tli-2026-v5.0.
As of 2026-09-06, GPT-6 Astra leads The Latent Score on the point estimate at 132.9 points among 19 ranked models. Rankings describe published capability evidence, not a statistically proven winner or the best model for every workload.
This page preserves a specific publication. Its data does not change when a new release is published.
Published rankings
| Rank | Model and evidence | Score (points) | Stability range (points) | Contributing results |
|---|---|---|---|---|
| 1 | GPT-6 Astra | 132.9 | 107.3–158.3 | 22 |
| 2 | Fable 5.1 | 131.0 | 101.4–157.5 | 20 |
| 3 | Opus 5 | 125.7 | 103.4–152.4 | 33 |
| 4 | Fable 5 | 117.1 | 92.8–141.4 | 30 |
| 5 | Grok 4.6 | 114.0 | 82–141.3 | 10 |
| 6 | Kimi K3 | 105.8 | 78.6–132 | 25 |
| 7 | GPT-5.6 Sol | 104.9 | 81.1–128 | 31 |
| 8 | Gemini 3.7 Flash | 104.6 | 81.1–131.6 | 28 |
| 9 | Gemini 3.8 Flash | 104.0 | 78–131.3 | 26 |
| 10 | Muse Spark 1.3 | 104.0 | 73.8–133.2 | 8 |
| 11 | GLM-5.3 | 94.9 | 68.9–120.6 | 23 |
| 12 | Qwen3.8-Max | 93.9 | 67.8–118.5 | 29 |
| 13 | GPT-5.6 Terra | 91.7 | 63.6–119.8 | 16 |
| 14 | Muse Spark 1.2 | 91.3 | 66.9–118.4 | 26 |
| 15 | DeepSeek V4 Pro | 88.8 | 62–115.5 | 23 |
| 16 | Sonnet 5 | 87.7 | 64–112.9 | 28 |
| 17 | GPT-5.6 Luna | 87.0 | 61.7–112.7 | 21 |
| 18 | DeepSeek V4 Flash | 84.1 | 53.2–120.6 | 5 |
| 19 | Qwen3.8-27B | 69.7 | 44.7–100.1 | 24 |
Methodology and limitations in this release
The scale is anchored to mean 100, SD 15 across 18 frozen reference models. It is not human IQ or a percentage. 41 benchmarks contribute across 7 domains.
One capability score combines evidence across seven equally weighted domains. Related benchmark variants share a family budget, and influence from one evaluator is capped.
Models are hard-ranked by the displayed point score. Ties use a stable model-name order. Intervals do not determine rank.
The interval reflects limited evidence, missing domains, correlated benchmark families and frozen calibration uncertainty under the scoring assumptions. It is not a promise that future scores will fall inside it 95% of the time. New conflicting evidence can widen it.
Each of seven capability domains has an equal, fixed importance budget. Missing domain abilities are conditional estimates, explicitly marked as unmeasured; individual missing benchmark results are never invented.
Selective reporting is assessed separately by withholding favorable and unfavorable results from reference models. The resulting sensitivity range has no probability interpretation and is not added to the main interval.
A model’s score and interval depend only on its own accepted evidence and the frozen methodology. Other models’ additions and updates cannot change them. A methodology revision receives a new version and can change every score.
Validation includes provider and benchmark-family holdouts, selective-reporting stress tests and replay of actual repository snapshots. Real launch history is limited and cannot establish universal future 95% coverage. This release checked 401 withheld benchmark results; normalized prediction error was 6.2% lower than a refitted legacy baseline on the same reviewed evidence. Only 2 first-recorded model arrivals across 2 providers had usable before/after evidence. Those repository dates are not verified launch dates. They do not validate launch-day coverage for every model, including Astra.
Reading this release
Compare model scores alongside their stability ranges and domain coverage. A higher point estimate does not establish superiority for every task. The linked model records disclose contributing and excluded results and their published configurations.
This release uses methodology version tli-2026-v5.0. Comparisons with another release should check both its methodology version and evidence date; a new data release does not necessarily mean the scoring method changed.
Citation and downloads
The Latent. “The Latent Score.” Release tli-2026-v5.0-2026-09-06, dataset dated 2026-09-06. Permanent release page.
Download public scores with definitions (JSON) · Data usage terms