The Latent Score: 2026-09-10

Release tli-2026-v1.0-2026-09-10-deepseek-3. Methodology tli-2026-v1.0.

As of 2026-09-10, GPT-6 Astra leads The Latent Score on the point estimate at 139.2 points among 20 ranked models. Rankings describe published capability evidence, not a statistically proven winner or the best model for every workload.

This page preserves a specific publication. Its data does not change when a new release is published.

Published rankings

RankModel and evidenceScore (points)Stability range (points)Contributing results
1GPT-6 Astra139.2128–150.254
2Fable 5.1134.7125.6–143.840
3Opus 5129.2123.6–134.869
4Fable 5120.0112.9–126.857
5Muse Spark 1.3114.2102.3–125.813
6DeepSeek V4.1-Flash113.099.7–128.310
7GPT-5.6 Sol111.9103.8–119.688
8Grok 4.6105.797.1–115.527
9Kimi K3104.095.2–112.567
10GLM-5.3103.892.2–115.342
11Gemini 3.8 Flash102.890.4–113.440
12Gemini 3.7 Flash101.491.2–111.541
13GPT-5.6 Terra95.188.8–101.354
14Qwen3.8-Max93.183.8–103.246
15Muse Spark 1.292.283.7–100.736
16DeepSeek V4 Pro86.277.1–95.338
17GPT-5.6 Luna81.373–90.154
18Sonnet 581.374.3–88.551
19Qwen3.8-27B70.561.5–79.340
20DeepSeek V4 Flash67.159.5–73.339

Methodology and limitations in this release

The scale is anchored to mean 100, SD 15 across 18 frozen reference models. It is not human IQ or a percentage. 92 benchmarks contribute across 7 domains.

One score summarizes demonstrated capability across published reasoning settings and seven equally weighted domains. Related benchmark variants share a family budget, and influence from one evaluator is capped.

Models are ranked by the displayed point score once they have four contributing benchmarks across two domains. Ties use stable model-name ordering. Ranks are a current ordering, not a confidence test. Scores are published immediately without a provisional label; early scores can move substantially as evidence arrives. Contributing benchmark counts and measured domain coverage appear beside each score.

The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.

Each of seven capability domains has an equal, fixed importance budget. These weights define the breadth valued by the index; they are not objectively optimal or learned importance weights. Missing domain abilities are conditional estimates, explicitly marked as unmeasured; individual missing benchmark results are never invented. Alternative weights and domain groupings are sensitivity checks, not replacements chosen to favor a model.

The separate reporting sensitivity range tests what happens when only the strongest or weakest standardized benchmark results are retained. It has no probability interpretation and is not combined with the score stability range. Stress-test coverage describes recovery of fuller-data scores under these deliberate masks, not real-world coverage or true intelligence. This test does not quantify the separate advantage from trying more reasoning settings.

This release checks 35 alternatives: halving or doubling each domain’s weight, and combining every pair of domains into one equally budgeted group. The largest rank movement relative to the equal-weight diagnostic is 2 places. These checks hold the published domain estimates fixed; they do not reassign benchmarks or refit the model. Rounded domain values can differ from published tie ordering. These are sensitivity scenarios, not confidence ranges or a search for preferred weights. Published scores and ranks are unchanged.

A model’s score and interval depend only on its own accepted evidence and the frozen methodology. Other models’ additions and updates cannot change them. A methodology revision receives a new version and can change every score.

Validation includes provider and benchmark-family holdouts, retrospective sparse-to-full score recovery and selective-reporting stress tests. These test prediction and score stability within the benchmark evidence; they do not establish capability on independent real-world tasks. Prospective tracking records dated evidence snapshots for later comparison. Existing models at tracking start are baselines, not observed launches. Independent-task evaluation requires a registered protocol and outcomes excluded from scoring; no completed independent-task validation is claimed.

Reading this release

Compare model scores alongside their stability ranges and domain coverage. A higher point estimate does not establish superiority for every task. The linked model records disclose contributing and excluded results and their published configurations.

This release uses methodology version tli-2026-v1.0. Comparisons with another release should check both its methodology version and evidence date; a new data release does not necessarily mean the scoring method changed.

Citation and downloads

The Latent. “The Latent Score.” Release tli-2026-v1.0-2026-09-10-deepseek-3, dataset dated 2026-09-10. Permanent release page.

Download public scores with definitions (JSON) · Data usage terms