Headline
The Latent Index 2026
The dot marks the point estimate; the bar spans the 95% bootstrap interval.
Ranks follow Latent score, highest first (1–15). Overlapping intervals mean two models can still be statistically close even when ordered by the point estimate.
- 1125.9
Latent score 125.9, 95% interval 119.8–132.1
- 95% interval
- 119.8–132.1
- Benchmarks
- 59
- Domains
- Knowledge & reasoning: +1.50 SDCoding & software engineering: +1.29 SD · leads this domainAgentic tool & computer use: +1.57 SD · leads this domainProfessional & real-world work: +2.17 SD · leads this domainMultimodal & vision: +1.29 SDCybersecurity: +0.71 SDPreference & communication: +1.06 SD
Audit
- Point estimate
- 125.9
- Score rank
- 1 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 1–2: replicates separate Opus 5 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 2122.3
Latent score 122.3, 95% interval 116.9–126.3
- 95% interval
- 116.9–126.3
- Benchmarks
- 72
- Domains
- Knowledge & reasoning: +1.72 SD · leads this domainCoding & software engineering: +1.28 SDAgentic tool & computer use: +1.08 SDProfessional & real-world work: +1.31 SDMultimodal & vision: +0.58 SDCybersecurity: +0.91 SD · leads this domainPreference & communication: +1.12 SD · leads this domain
Audit
- Point estimate
- 122.3
- Score rank
- 2 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 1–2: replicates separate Fable 5 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 3111.5
Latent score 111.5, 95% interval 106.1–116.6
- 95% interval
- 106.1–116.6
- Benchmarks
- 72
- Domains
- Knowledge & reasoning: +1.01 SDCoding & software engineering: +0.90 SDAgentic tool & computer use: +0.74 SDProfessional & real-world work: +0.20 SDMultimodal & vision: +0.66 SDCybersecurity: +0.87 SDPreference & communication: +0.22 SD
Audit
- Point estimate
- 111.5
- Score rank
- 3 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 3–6: replicates separate GPT-5.6 Sol from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 4109.5
Latent score 109.5, 95% interval 98.7–115.7
- 95% interval
- 98.7–115.7
- Benchmarks
- 31
- Domains
- Knowledge & reasoning: +0.25 SDCoding & software engineering: -0.29 SDAgentic tool & computer use: +1.47 SDProfessional & real-world work: +0.39 SDMultimodal & vision: +0.02 SDCybersecurity: not measuredPreference & communication: +0.93 SD
Audit
- Point estimate
- 109.5
- Score rank
- 4 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 3–8: replicates separate Muse Spark 1.2 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 5108.7
Latent score 108.7, 95% interval 102.7–112.8
- 95% interval
- 102.7–112.8
- Benchmarks
- 63
- Domains
- Knowledge & reasoning: +0.05 SDCoding & software engineering: +0.35 SDAgentic tool & computer use: +0.63 SDProfessional & real-world work: +0.57 SDMultimodal & vision: +0.47 SDCybersecurity: -0.10 SDPreference & communication: +1.04 SD
Audit
- Point estimate
- 108.7
- Score rank
- 5 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 3–8: replicates separate Kimi K3 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 6107.7
Latent score 107.7, 95% interval 99.7–113.9
- 95% interval
- 99.7–113.9
- Benchmarks
- 23
- Domains
- Knowledge & reasoning: -0.25 SDCoding & software engineering: +0.90 SDAgentic tool & computer use: +0.44 SDProfessional & real-world work: +0.99 SDMultimodal & vision: not measuredCybersecurity: +0.39 SDPreference & communication: not measured
Audit
- Point estimate
- 107.7
- Score rank
- 6 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 3–8: replicates separate GLM-5.3 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 7103.4
Latent score 103.4, 95% interval 98.2–111.4
- 95% interval
- 98.2–111.4
- Benchmarks
- 37
- Domains
- Knowledge & reasoning: +0.30 SDCoding & software engineering: +0.79 SDAgentic tool & computer use: +0.54 SDProfessional & real-world work: +0.47 SDMultimodal & vision: -2.55 SDCybersecurity: +0.01 SDPreference & communication: -0.58 SD
Audit
- Point estimate
- 103.4
- Score rank
- 7 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 4–9: replicates separate Grok 4.6 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 8101.7
Latent score 101.7, 95% interval 95.8–106.2
- 95% interval
- 95.8–106.2
- Benchmarks
- 51
- Domains
- Knowledge & reasoning: +0.07 SDCoding & software engineering: -0.36 SDAgentic tool & computer use: -0.37 SDProfessional & real-world work: +0.18 SDMultimodal & vision: +1.32 SD · leads this domainCybersecurity: -0.19 SDPreference & communication: +0.74 SD
Audit
- Point estimate
- 101.7
- Score rank
- 8 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 7–10: replicates separate Qwen3.8-Max from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 998.4
Latent score 98.4, 95% interval 87–110.8
- 95% interval
- 87–110.8
- Benchmarks
- 22
- Domains
- Knowledge & reasoning: +0.81 SDCoding & software engineering: +0.10 SDAgentic tool & computer use: -1.09 SDProfessional & real-world work: -1.32 SDMultimodal & vision: -0.52 SDCybersecurity: not measuredPreference & communication: +0.71 SD
Audit
- Point estimate
- 98.4
- Score rank
- 9 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 4–11: replicates separate Gemini 3.7 Flash from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 1096.1
Latent score 96.1, 95% interval 91.7–101.8
- 95% interval
- 91.7–101.8
- Benchmarks
- 48
- Domains
- Knowledge & reasoning: -0.26 SDCoding & software engineering: +0.25 SDAgentic tool & computer use: -0.21 SDProfessional & real-world work: -0.56 SDMultimodal & vision: +0.10 SDCybersecurity: +0.33 SDPreference & communication: -0.62 SD
Audit
- Point estimate
- 96.1
- Score rank
- 10 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 8–11: replicates separate GPT-5.6 Terra from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 1192.1
Latent score 92.1, 95% interval 85.8–100.8
- 95% interval
- 85.8–100.8
- Benchmarks
- 35
- Domains
- Knowledge & reasoning: -0.31 SDCoding & software engineering: -1.09 SDAgentic tool & computer use: -0.09 SDProfessional & real-world work: -0.81 SDMultimodal & vision: not measuredCybersecurity: +0.59 SDPreference & communication: -0.53 SD
Audit
- Point estimate
- 92.1
- Score rank
- 11 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 9–13: replicates separate DeepSeek V4 Pro from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 1289.8
Latent score 89.8, 95% interval 80.5–95.3
- 95% interval
- 80.5–95.3
- Benchmarks
- 47
- Domains
- Knowledge & reasoning: -0.42 SDCoding & software engineering: -0.19 SDAgentic tool & computer use: -0.82 SDProfessional & real-world work: -0.25 SDMultimodal & vision: -0.66 SDCybersecurity: -3.01 SDPreference & communication: -0.75 SD
Audit
- Point estimate
- 89.8
- Score rank
- 12 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 11–13: replicates separate Sonnet 5 from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 1388.0
Latent score 88.0, 95% interval 83.2–95
- 95% interval
- 83.2–95
- Benchmarks
- 47
- Domains
- Knowledge & reasoning: -0.93 SDCoding & software engineering: -0.10 SDAgentic tool & computer use: -0.70 SDProfessional & real-world work: -0.77 SDMultimodal & vision: -0.57 SDCybersecurity: -0.09 SDPreference & communication: -1.18 SD
Audit
- Point estimate
- 88.0
- Score rank
- 13 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 11–13: replicates separate GPT-5.6 Luna from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 14
DeepSeek V4 Flash
DeepSeek
73.0Latent score 73.0, 95% interval 67.3–81.2
- 95% interval
- 67.3–81.2
- Benchmarks
- 42
- Domains
- Knowledge & reasoning: -1.31 SDCoding & software engineering: -1.45 SDAgentic tool & computer use: -1.90 SDProfessional & real-world work: -1.57 SDMultimodal & vision: not measuredCybersecurity: -0.41 SDPreference & communication: -2.14 SD
Audit
- Point estimate
- 73.0
- Score rank
- 14 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 14–15: replicates separate DeepSeek V4 Flash from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database → - 1572.0
Latent score 72.0, 95% interval 62.7–85.9
- 95% interval
- 62.7–85.9
- Benchmarks
- 21
- Domains
- Knowledge & reasoning: -2.23 SDCoding & software engineering: -2.39 SDAgentic tool & computer use: -1.28 SDProfessional & real-world work: -0.99 SDMultimodal & vision: -0.14 SDCybersecurity: not measuredPreference & communication: not measured
Audit
- Point estimate
- 72.0
- Score rank
- 15 of 15 by Latent point estimate
- Bootstrap separation
- Bootstrap separation 14–15: replicates separate Qwen3.8-27B from every model outside this range at ≥90% probability of superiority, but cannot order it within the range.
- Uncertainty
- 95% bootstrap interval from stratified domain resampling
- Methodology
- tli-2026-v3.4
- Dataset
- 2026-08-23
Open the benchmark database →