The Latent Score: 2026-09-06
Release tli-2026-v1.0-2026-09-06-presentation-1. Methodology tli-2026-v1.0.
As of 2026-09-06, GPT-6 Astra leads The Latent Score on the point estimate at 139.0 points among 19 ranked models. Rankings describe published capability evidence, not a statistically proven winner or the best model for every workload.
This page preserves a specific publication. Its data does not change when a new release is published.
Published rankings
| Rank | Model and evidence | Score (points) | Stability range (points) | Contributing results |
|---|---|---|---|---|
| 1 | GPT-6 Astra | 139.0 | 128.8–149.3 | 45 |
| 2 | Fable 5.1 | 131.2 | 123.3–138.1 | 35 |
| 3 | Opus 5 | 129.5 | 123.8–135.1 | 68 |
| 4 | Fable 5 | 119.6 | 112.5–126.6 | 56 |
| 5 | Muse Spark 1.3 | 114.2 | 102.3–125.8 | 13 |
| 6 | GPT-5.6 Sol | 111.3 | 103.2–119.2 | 86 |
| 7 | Grok 4.6 | 105.7 | 97.1–115.5 | 27 |
| 8 | Kimi K3 | 104.9 | 95.4–113.7 | 64 |
| 9 | Gemini 3.8 Flash | 102.8 | 90.4–113.4 | 40 |
| 10 | GLM-5.3 | 101.6 | 90.8–113.9 | 39 |
| 11 | Gemini 3.7 Flash | 101.4 | 91.2–111.5 | 41 |
| 12 | GPT-5.6 Terra | 95.1 | 88.8–101.3 | 54 |
| 13 | Qwen3.8-Max | 94.6 | 84.4–105 | 43 |
| 14 | Muse Spark 1.2 | 92.2 | 83.7–100.7 | 36 |
| 15 | DeepSeek V4 Pro | 83.2 | 74.1–93.9 | 35 |
| 16 | GPT-5.6 Luna | 81.3 | 73–90.1 | 54 |
| 17 | Sonnet 5 | 81.3 | 74.3–88.5 | 51 |
| 18 | Qwen3.8-27B | 71.0 | 61.8–79.9 | 39 |
| 19 | DeepSeek V4 Flash | 67.2 | 59.6–74 | 37 |
Methodology and limitations in this release
The scale is anchored to mean 100, SD 15 across 18 frozen reference models. It is not human IQ or a percentage. 89 benchmarks contribute across 7 domains.
One score summarizes demonstrated capability across published reasoning settings and seven equally weighted domains. Related benchmark variants share a family budget, and influence from one evaluator is capped.
Models are ranked by the displayed point score once they have four contributing benchmarks across two domains. Ties use stable model-name ordering. Ranks are a current ordering, not a confidence test. Scores are published immediately without a provisional label; early scores can move substantially as evidence arrives. Contributing benchmark counts and measured domain coverage appear beside each score.
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
Each of seven capability domains has an equal, fixed importance budget. These weights define the breadth valued by the index; they are not objectively optimal or learned importance weights. Missing domain abilities are conditional estimates, explicitly marked as unmeasured; individual missing benchmark results are never invented. Alternative weights and domain groupings are sensitivity checks, not replacements chosen to favor a model.
The separate reporting sensitivity range tests what happens when only the strongest or weakest standardized benchmark results are retained. It has no probability interpretation and is not combined with the score stability range. Stress-test coverage describes recovery of fuller-data scores under these deliberate masks, not real-world coverage or true intelligence. This test does not quantify the separate advantage from trying more reasoning settings.
This release checks 35 alternatives: halving or doubling each domain’s weight, and combining every pair of domains into one equally budgeted group. The largest rank movement relative to the equal-weight diagnostic is 2 places. These checks hold the published domain estimates fixed; they do not reassign benchmarks or refit the model. Rounded domain values can differ from published tie ordering. These are sensitivity scenarios, not confidence ranges or a search for preferred weights. Published scores and ranks are unchanged.
A model’s score and interval depend only on its own accepted evidence and the frozen methodology. Other models’ additions and updates cannot change them. A methodology revision receives a new version and can change every score.
Validation includes provider and benchmark-family holdouts, retrospective sparse-to-full score recovery and selective-reporting stress tests. These test prediction and score stability within the benchmark evidence; they do not establish capability on independent real-world tasks. Prospective tracking records dated evidence snapshots for later comparison. Existing models at tracking start are baselines, not observed launches. Independent-task evaluation requires a registered protocol and outcomes excluded from scoring; no completed independent-task validation is claimed. This release checked 763 withheld benchmark results; normalized prediction error was 4.4% lower than a refitted legacy baseline on the same reviewed evidence. Only 4 first-recorded model arrivals across 3 providers had usable before/after evidence. Those repository dates are not verified launch dates. They do not validate launch-day coverage for every model, including Astra. The main score interval recovered fuller-data scores in 52 of 54 random held-out evidence cases. Selective favorable and unfavorable publication had only 13% and 13% recovery of fuller-data scores in the main bar under these deliberate stress masks; this is not an estimate of real-world coverage. Reporting sensitivity is shown separately without a probability claim. 4 historical arrivals required a conservative fallback because their earlier reference data were insufficient to train the correction. Those replays do not validate the tighter empirical bands. Calibration needs reference models from at least three providers; a narrow bar is never forced by a width cap.
Reading this release
Compare model scores alongside their stability ranges and domain coverage. A higher point estimate does not establish superiority for every task. The linked model records disclose contributing and excluded results and their published configurations.
This release uses methodology version tli-2026-v1.0. Comparisons with another release should check both its methodology version and evidence date; a new data release does not necessarily mean the scoring method changed.
Citation and downloads
The Latent. “The Latent Score.” Release tli-2026-v1.0-2026-09-06-presentation-1, dataset dated 2026-09-06. Permanent release page.
Download public scores with definitions (JSON) · Data usage terms