The Latent Score

One number for how smart each frontier model actually is.

The Latent Score is the general-ability estimate from a hierarchical latent factor model fitted to 73 published benchmark results across 7 capability domains. The Latent does not run the evaluations; it estimates the single ability that best explains all of them at once.

The scale is anchored at mean 100, SD 15 across the current cohort; higher is smarter. Every score carries a 95% bootstrap interval; models still waiting on benchmark coverage are labeled provisional and stay outside the official ranking.

125.9leading score · Opus 5
73benchmarks inside the fit
7capability domains
2026-V3.4frozen methodology version

Headline

The Latent Index 2026

The dot marks the point estimate; the bar spans the 95% bootstrap interval.
Ranks follow Latent score, highest first (1–15). Overlapping intervals mean two models can still be statistically close even when ordered by the point estimate.

  1. 1

    Opus 5

    Anthropic

    125.9
    Latent score 125.9, 95% interval 119.8132.1
    95% interval
    119.8132.1
    Benchmarks
    59
    Domains
    Knowledge & reasoning: +1.50 SDCoding & software engineering: +1.29 SD · leads this domainAgentic tool & computer use: +1.57 SD · leads this domainProfessional & real-world work: +2.17 SD · leads this domainMultimodal & vision: +1.29 SDCybersecurity: +0.71 SDPreference & communication: +1.06 SD
    Audit
    Point estimate
    125.9
    Score rank
    1 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 1–2: replicates separate Opus 5 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  2. 2

    Fable 5

    Anthropic

    122.3
    Latent score 122.3, 95% interval 116.9126.3
    95% interval
    116.9126.3
    Benchmarks
    72
    Domains
    Knowledge & reasoning: +1.72 SD · leads this domainCoding & software engineering: +1.28 SDAgentic tool & computer use: +1.08 SDProfessional & real-world work: +1.31 SDMultimodal & vision: +0.58 SDCybersecurity: +0.91 SD · leads this domainPreference & communication: +1.12 SD · leads this domain
    Audit
    Point estimate
    122.3
    Score rank
    2 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 1–2: replicates separate Fable 5 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  3. 3

    GPT-5.6 Sol

    OpenAI

    111.5
    Latent score 111.5, 95% interval 106.1116.6
    95% interval
    106.1116.6
    Benchmarks
    72
    Domains
    Knowledge & reasoning: +1.01 SDCoding & software engineering: +0.90 SDAgentic tool & computer use: +0.74 SDProfessional & real-world work: +0.20 SDMultimodal & vision: +0.66 SDCybersecurity: +0.87 SDPreference & communication: +0.22 SD
    Audit
    Point estimate
    111.5
    Score rank
    3 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 3–6: replicates separate GPT-5.6 Sol from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  4. 4

    Muse Spark 1.2

    Meta

    109.5
    Latent score 109.5, 95% interval 98.7115.7
    95% interval
    98.7115.7
    Benchmarks
    31
    Domains
    Knowledge & reasoning: +0.25 SDCoding & software engineering: -0.29 SDAgentic tool & computer use: +1.47 SDProfessional & real-world work: +0.39 SDMultimodal & vision: +0.02 SDCybersecurity: not measuredPreference & communication: +0.93 SD
    Audit
    Point estimate
    109.5
    Score rank
    4 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 3–8: replicates separate Muse Spark 1.2 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  5. 5

    Kimi K3

    Moonshot AI

    108.7
    Latent score 108.7, 95% interval 102.7112.8
    95% interval
    102.7112.8
    Benchmarks
    63
    Domains
    Knowledge & reasoning: +0.05 SDCoding & software engineering: +0.35 SDAgentic tool & computer use: +0.63 SDProfessional & real-world work: +0.57 SDMultimodal & vision: +0.47 SDCybersecurity: -0.10 SDPreference & communication: +1.04 SD
    Audit
    Point estimate
    108.7
    Score rank
    5 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 3–8: replicates separate Kimi K3 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  6. 6

    GLM-5.3

    Z.ai

    107.7
    Latent score 107.7, 95% interval 99.7113.9
    95% interval
    99.7113.9
    Benchmarks
    23
    Domains
    Knowledge & reasoning: -0.25 SDCoding & software engineering: +0.90 SDAgentic tool & computer use: +0.44 SDProfessional & real-world work: +0.99 SDMultimodal & vision: not measuredCybersecurity: +0.39 SDPreference & communication: not measured
    Audit
    Point estimate
    107.7
    Score rank
    6 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 3–8: replicates separate GLM-5.3 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  7. 7

    Grok 4.6

    xAI

    103.4
    Latent score 103.4, 95% interval 98.2111.4
    95% interval
    98.2111.4
    Benchmarks
    37
    Domains
    Knowledge & reasoning: +0.30 SDCoding & software engineering: +0.79 SDAgentic tool & computer use: +0.54 SDProfessional & real-world work: +0.47 SDMultimodal & vision: -2.55 SDCybersecurity: +0.01 SDPreference & communication: -0.58 SD
    Audit
    Point estimate
    103.4
    Score rank
    7 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 4–9: replicates separate Grok 4.6 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  8. 8

    Qwen3.8-Max

    Alibaba

    101.7
    Latent score 101.7, 95% interval 95.8106.2
    95% interval
    95.8106.2
    Benchmarks
    51
    Domains
    Knowledge & reasoning: +0.07 SDCoding & software engineering: -0.36 SDAgentic tool & computer use: -0.37 SDProfessional & real-world work: +0.18 SDMultimodal & vision: +1.32 SD · leads this domainCybersecurity: -0.19 SDPreference & communication: +0.74 SD
    Audit
    Point estimate
    101.7
    Score rank
    8 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 7–10: replicates separate Qwen3.8-Max from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  9. 9

    Gemini 3.7 Flash

    Google

    98.4
    Latent score 98.4, 95% interval 87110.8
    95% interval
    87110.8
    Benchmarks
    22
    Domains
    Knowledge & reasoning: +0.81 SDCoding & software engineering: +0.10 SDAgentic tool & computer use: -1.09 SDProfessional & real-world work: -1.32 SDMultimodal & vision: -0.52 SDCybersecurity: not measuredPreference & communication: +0.71 SD
    Audit
    Point estimate
    98.4
    Score rank
    9 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 4–11: replicates separate Gemini 3.7 Flash from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  10. 10

    GPT-5.6 Terra

    OpenAI

    96.1
    Latent score 96.1, 95% interval 91.7101.8
    95% interval
    91.7101.8
    Benchmarks
    48
    Domains
    Knowledge & reasoning: -0.26 SDCoding & software engineering: +0.25 SDAgentic tool & computer use: -0.21 SDProfessional & real-world work: -0.56 SDMultimodal & vision: +0.10 SDCybersecurity: +0.33 SDPreference & communication: -0.62 SD
    Audit
    Point estimate
    96.1
    Score rank
    10 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 8–11: replicates separate GPT-5.6 Terra from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  11. 11

    DeepSeek V4 Pro

    DeepSeek

    92.1
    Latent score 92.1, 95% interval 85.8100.8
    95% interval
    85.8100.8
    Benchmarks
    35
    Domains
    Knowledge & reasoning: -0.31 SDCoding & software engineering: -1.09 SDAgentic tool & computer use: -0.09 SDProfessional & real-world work: -0.81 SDMultimodal & vision: not measuredCybersecurity: +0.59 SDPreference & communication: -0.53 SD
    Audit
    Point estimate
    92.1
    Score rank
    11 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 9–13: replicates separate DeepSeek V4 Pro from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  12. 12

    Sonnet 5

    Anthropic

    89.8
    Latent score 89.8, 95% interval 80.595.3
    95% interval
    80.595.3
    Benchmarks
    47
    Domains
    Knowledge & reasoning: -0.42 SDCoding & software engineering: -0.19 SDAgentic tool & computer use: -0.82 SDProfessional & real-world work: -0.25 SDMultimodal & vision: -0.66 SDCybersecurity: -3.01 SDPreference & communication: -0.75 SD
    Audit
    Point estimate
    89.8
    Score rank
    12 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 11–13: replicates separate Sonnet 5 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  13. 13

    GPT-5.6 Luna

    OpenAI

    88.0
    Latent score 88.0, 95% interval 83.295
    95% interval
    83.295
    Benchmarks
    47
    Domains
    Knowledge & reasoning: -0.93 SDCoding & software engineering: -0.10 SDAgentic tool & computer use: -0.70 SDProfessional & real-world work: -0.77 SDMultimodal & vision: -0.57 SDCybersecurity: -0.09 SDPreference & communication: -1.18 SD
    Audit
    Point estimate
    88.0
    Score rank
    13 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 11–13: replicates separate GPT-5.6 Luna from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  14. 14

    DeepSeek V4 Flash

    DeepSeek

    73.0
    Latent score 73.0, 95% interval 67.381.2
    95% interval
    67.381.2
    Benchmarks
    42
    Domains
    Knowledge & reasoning: -1.31 SDCoding & software engineering: -1.45 SDAgentic tool & computer use: -1.90 SDProfessional & real-world work: -1.57 SDMultimodal & vision: not measuredCybersecurity: -0.41 SDPreference & communication: -2.14 SD
    Audit
    Point estimate
    73.0
    Score rank
    14 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 14–15: replicates separate DeepSeek V4 Flash from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →
  15. 15

    Qwen3.8-27B

    Alibaba

    72.0
    Latent score 72.0, 95% interval 62.785.9
    95% interval
    62.785.9
    Benchmarks
    21
    Domains
    Knowledge & reasoning: -2.23 SDCoding & software engineering: -2.39 SDAgentic tool & computer use: -1.28 SDProfessional & real-world work: -0.99 SDMultimodal & vision: -0.14 SDCybersecurity: not measuredPreference & communication: not measured
    Audit
    Point estimate
    72.0
    Score rank
    15 of 15 by Latent point estimate
    Bootstrap separation
    Bootstrap separation 14–15: replicates separate Qwen3.8-27B from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
    Uncertainty
    95% bootstrap interval from stratified domain resampling
    Methodology
    tli-2026-v3.4
    Dataset
    2026-08-23
    Open the benchmark database →

Domain profile

Seven abilities, one general factor

Domain scores are in SD units within each domain.
Browse the benchmark database →

Knowledge & reasoning

11 benchmarks in this domain
  1. Fable 5+1.72
  2. Opus 5+1.50
  3. GPT-5.6 Sol+1.01
  4. Gemini 3.7 Flash+0.81
  5. Grok 4.6+0.30
  6. Muse Spark 1.2+0.25
  7. Qwen3.8-Max+0.07
  8. Kimi K3+0.05
  9. GLM-5.3-0.25
  10. GPT-5.6 Terra-0.26
  11. DeepSeek V4 Pro-0.31
  12. Sonnet 5-0.42
  13. GPT-5.6 Luna-0.93
  14. DeepSeek V4 Flash-1.31
  15. Qwen3.8-27B-2.23

Coding & software engineering

19 benchmarks in this domain
  1. Opus 5+1.29
  2. Fable 5+1.28
  3. GPT-5.6 Sol+0.90
  4. GLM-5.3+0.90
  5. Grok 4.6+0.79
  6. Kimi K3+0.35
  7. GPT-5.6 Terra+0.25
  8. Gemini 3.7 Flash+0.10
  9. GPT-5.6 Luna-0.10
  10. Sonnet 5-0.19
  11. Muse Spark 1.2-0.29
  12. Qwen3.8-Max-0.36
  13. DeepSeek V4 Pro-1.09
  14. DeepSeek V4 Flash-1.45
  15. Qwen3.8-27B-2.39

Agentic tool & computer use

13 benchmarks in this domain
  1. Opus 5+1.57
  2. Muse Spark 1.2+1.47
  3. Fable 5+1.08
  4. GPT-5.6 Sol+0.74
  5. Kimi K3+0.63
  6. Grok 4.6+0.54
  7. GLM-5.3+0.44
  8. DeepSeek V4 Pro-0.09
  9. GPT-5.6 Terra-0.21
  10. Qwen3.8-Max-0.37
  11. GPT-5.6 Luna-0.70
  12. Sonnet 5-0.82
  13. Gemini 3.7 Flash-1.09
  14. Qwen3.8-27B-1.28
  15. DeepSeek V4 Flash-1.90

Professional & real-world work

17 benchmarks in this domain
  1. Opus 5+2.17
  2. Fable 5+1.31
  3. GLM-5.3+0.99
  4. Kimi K3+0.57
  5. Grok 4.6+0.47
  6. Muse Spark 1.2+0.39
  7. GPT-5.6 Sol+0.20
  8. Qwen3.8-Max+0.18
  9. Sonnet 5-0.25
  10. GPT-5.6 Terra-0.56
  11. GPT-5.6 Luna-0.77
  12. DeepSeek V4 Pro-0.81
  13. Qwen3.8-27B-0.99
  14. Gemini 3.7 Flash-1.32
  15. DeepSeek V4 Flash-1.57

Multimodal & vision

7 benchmarks in this domain
  1. Qwen3.8-Max+1.32
  2. Opus 5+1.29
  3. GPT-5.6 Sol+0.66
  4. Fable 5+0.58
  5. Kimi K3+0.47
  6. GPT-5.6 Terra+0.10
  7. Muse Spark 1.2+0.02
  8. Qwen3.8-27B-0.14
  9. Gemini 3.7 Flash-0.52
  10. GPT-5.6 Luna-0.57
  11. Sonnet 5-0.66
  12. Grok 4.6-2.55

Cybersecurity

2 benchmarks in this domain
  1. Fable 5+0.91
  2. GPT-5.6 Sol+0.87
  3. Opus 5+0.71
  4. DeepSeek V4 Pro+0.59
  5. GLM-5.3+0.39
  6. GPT-5.6 Terra+0.33
  7. Grok 4.6+0.01
  8. GPT-5.6 Luna-0.09
  9. Kimi K3-0.10
  10. Qwen3.8-Max-0.19
  11. DeepSeek V4 Flash-0.41
  12. Sonnet 5-3.01

Preference & communication

4 benchmarks in this domain
  1. Fable 5+1.12
  2. Opus 5+1.06
  3. Kimi K3+1.04
  4. Muse Spark 1.2+0.93
  5. Qwen3.8-Max+0.74
  6. Gemini 3.7 Flash+0.71
  7. GPT-5.6 Sol+0.22
  8. DeepSeek V4 Pro-0.53
  9. Grok 4.6-0.58
  10. GPT-5.6 Terra-0.62
  11. Sonnet 5-0.75
  12. GPT-5.6 Luna-1.18
  13. DeepSeek V4 Flash-2.14

Frozen until the next major version

How the score is estimated

Full methodology →

Published evidence only

The Latent does not run evaluations. It estimates one general ability from published benchmark results, comparing models only on the rows they both have. Missing cells are never imputed.

Hierarchical by design

Overlapping benchmarks are grouped into 7 capability domains so duplication does not dominate the headline number. Domain abilities are then combined into The Latent Score without editorial pillar weights.

Uncertainty, shown honestly

Every score ships with a 95% interval. Ranks are ranges: two models are only ordered when bootstrap probability of superiority clears a high bar. Thin coverage widens the interval or keeps a model provisional.

Checked, not copied

Composite meta-indices are held out of the fit and used as out-of-formula checks (Spearman 0.92 against the strongest held-out composite). Exact fitted weights and hyperparameters stay proprietary; the principles are on /methodology.

Cost, latency, throughput, and reliability are not inputs to The Latent Score. Cost enters only on Latent Value & Efficiency, where the workload, source, and penalty are explicit.