Comparable rows with enough coverage to support an estimate, each assigned to one researched capability domain.
Methodology · tli-2026-v3.4
Other organizations run the benchmarks. The Latent makes sense of them.
The Latent Score is a general-ability estimate for frontier models, built so that missing data, benchmark duplication, and false precision are handled in the open, while the fitted recipe that produces the number stays proprietary.
This page states the principles the index is frozen to. It is not a reimplementation guide. Exact hyperparameters, fitted loadings, and bootstrap artifacts are not published.
Layer one
The benchmark database is the evidence layer
The database holds every credible published result we track, not only the rows that enter the index. Each row keeps its native unit, coverage, and source caveats. The matrix is sparse, and the index is built for exactly that sparsity.
Composite meta-indices aggregate rows already in the sheet. Including them would double-count; they are used as out-of-formula checks instead.
Rows that cover too few models stay visible in the database and join the fit as coverage arrives, without rewriting the method.
Principles
What The Latent Score is designed to do
Glass-box principles · closed-box recipe
- 01Evidence, not evaluationsThe Latent does not run benchmarks. It estimates ability from published results that are version-pinned, metric-defined, and provenance-tagged.
- 02No editorial pillar weightsHow much each benchmark and domain can pull on the headline number is estimated from the data, not assigned by hand. Exact fitted weights stay proprietary.
- 03No imputationModels are compared only on benchmarks they both ran. Missing cells are never filled in; thin coverage shows up as wider intervals or provisional status.
- 04Deduplicate by structureOverlapping benchmarks are grouped into seven capability domains so ten coding rows sharpen one coding ability instead of counting ten times.
- 05One general factorDomain abilities feed a general-ability estimate, The Latent Score, on a cohort-anchored scale (mean 100, SD 15). Higher is smarter.
- 06Honest uncertaintyEvery score carries a 95% interval. Published ranks are ranges: two models are only ordered when probability of superiority clears a high bar.
Deduplication by structure
Seven capability domains
Domains organize evidence; they are not editorial score weights.
Eighteen coding benchmarks are not eighteen votes for coding. Within a domain, overlapping rows jointly sharpen one ability estimate. How strongly each domain then tracks the general factor is estimated from this cohort, and those loadings are not published as a public recipe.
Academic knowledge, mathematics, abstract reasoning, factuality, and long-context reasoning.
Agentic terminal work, issue resolution, long-horizon engineering, competitive coding, and autonomous ML engineering.
Tool orchestration, MCP servers, business-workflow automation, browsing, and GUI computer use.
Economically valuable knowledge work: legal, finance, tax, medical, enterprise documents, and occupational workflows.
Chart, image, video, and handwriting understanding.
Vulnerability analysis and exploit development capability.
Arena-style human preference, emotional intelligence, creative writing, and design preference.
Uncertainty
Intervals and rank ranges, not false precision
95% intervals come from resampling the evidence and refitting the index. Where the data cannot separate two models at high confidence, the board shows a rank range instead of inventing a strict order.
Models below a frozen coverage criterion are labeled provisional. Their results still inform the fit, but they hold no official rank until coverage fills in. Graduation is a data event, not a methodology rewrite.
Checks
Out-of-formula agreement, standing stress tests
The index is checked against composite meta-indices it never uses as inputs, and against stability and selective-reporting stress analyses that ship with every internal release. Public pages report the conclusions, not the harness.
- AA Intelligence Index v4.1.1 (score): Spearman 0.92 (n=15)
- Vals Index: Spearman 0.87 (n=13)
- Vals Multimodal Index: Spearman 0.78 (n=9)
- AIIQ Composite IQ: Spearman 0.66 (n=6)
The Latent Index 2026
Living data, frozen method
New published results extend the sheet and scores recompute under a frozen methodology version (tli-2026-v3.4). Numbers can move when the evidence grows; the method does not silently rewrite itself. Changing the estimator, domain taxonomy, or eligibility rules requires a new version string, never a quiet edit.
- Exact fitted benchmark and domain weights
- Hyperparameters, priors, and clipping constants
- Bootstrap configuration and full pairwise matrices
- Internal reference implementation and audit notebooks
Latent Efficiency · le-2026-v2.0 · Latent Value · lv-2026-v3.0
Cost is a separate, explicit layer
Cost never enters The Latent Score. The value page applies rule-based purchasing criteria — no scalar value score. A model that a cheaper-or-equal ranked model beats on capability is never recommended ahead of its dominator (point estimates decide placement; a dominance claim is only presented as decisive when the dominator's capability advantage holds in ≥90% of bootstrap replicates). Non-dominated models form a buying ladder, cheapest first, each rung stating what the extra spend buys; the headline Value Pick is the cheapest ladder model clearing a declared capability floor (Latent ≥ 80, a versioned policy on the cohort-anchored 100±15 scale — not an empirical discovery). The headline lens is Artificial Analysis's independently observed cost to complete one benchmark task (reasoning tokens included); the configuration board applies the same rules to every published reasoning effort of the ranked models; the secondary lens is a fixed listed-price basket (1,000,000 input + 250,000 output tokens at listed first-party rates). The scalar Efficiency and Value scores of the previous board versions remain in the published snapshots for reproducibility, but no displayed ordering uses them.
What never enters the score
Separate signals stay separate
- Cost, latency, throughput, and reliability
- Editorial taste and vendor marketing claims
- Imputed or extrapolated results for missing cells
- Composite indices built from rows already in the sheet
Reproducibility of the published number
Every headline score has an audit trail on the site
Open a score on The Latent Score to see its point estimate, 95% interval, benchmark count, and per-domain fingerprint. The underlying published cells remain on the benchmark database. The fit that turns those cells into The Latent Score is deterministic inside The Latent: same sheet snapshot, same methodology version, same number.
“According to The Latent Score” means a statistical aggregation of published benchmark results under a frozen proprietary methodology, not a claim that The Latent ran every evaluation.