Latent Value & Efficiency
Which model gives you the most intelligence for the money?
A rule-based buying guide: cheaper and smarter always wins, and the headline pick is the cheapest genuinely capable model. Cost is what a model actually spent per benchmark task, measured independently by Artificial Analysis — reasoning tokens included, and price never feeds back into the intelligence estimate.
Best bang for your buck
GPT-5.6 Luna
No cheaper model comes close; every smarter model costs multiples more.
- Latent Score
- 88.0
- Observed cost / task
- $0.047
- vs DeepSeek V4 Flash
- 57% cheaper
+15.0 Latent points
Capability × observed cost
The Latent Frontier
Up and to the left is better. Filled red points are the buying ladder: no cheaper option beats them. Every published reasoning configuration is plotted.
- Buying ladder
- Dominated, within uncertainty of the ladder
- Dominated
- Estimated (est.): derived from public effort benchmarks, never ranked
Qwen3.8-27B has no Artificial Analysis cost measurement and cannot be plotted. GLM-5.3 has no Artificial Analysis cost measurement; its point uses an estimated cost transferred from its predecessor's measured rate at price parity. Cost data: Artificial Analysis Intelligence Index v4.1.1, retrieved 2026-08-17.
Observed cost · Artificial Analysis
Spend more for more intelligence
No scalar value score. A model dominated by a cheaper-and-smarter model is never recommended ahead of it; non-dominated models form a buying ladder, cheapest first; the Value Pick is the cheapest ladder model clearing the declared Latent ≥ 80 floor.
GPT-5.6 Luna
OpenAI
Best bang for buck
- Latent Score
- 88.0
- Cost
- $0.047/task
Audit
- Measured configuration
- GPT-5.6 Luna (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 83.2 to 95.0
- Note
- AA now measures Luna at max effort: the configuration the sheet pins (retrieved 2026-08-17).
+$0.20/task · +4.1 Latent points · 5.3× the cost
DeepSeek V4 Pro
DeepSeek
- Latent Score
- 92.1
- Cost
- $0.25/task
Audit
- Measured configuration
- DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 85.8 to 100.8
- Note
- AA measures the 0813 API snapshot at Max effort: the configuration the sheet pins.
+$0.15/task · +17.4 Latent points · 1.6× the cost
Muse Spark 1.2
Meta
- Latent Score
- 109.5
- Cost
- $0.40/task
Audit
- Measured configuration
- Muse Spark 1.2 (xhigh)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 98.7 to 115.7
+$0.83/task · +2.0 Latent points · 3.1× the cost
GPT-5.6 Sol
OpenAI
- Latent Score
- 111.5
- Cost
- $1.23/task
Audit
- Measured configuration
- GPT-5.6 Sol (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 106.1 to 116.6
+$1.11/task · +14.4 Latent points · 1.9× the cost
Opus 5
Anthropic
Maximum intelligence
- Latent Score
- 125.9
- Cost
- $2.34/task
Audit
- Measured configuration
- Claude Opus 5 (Adaptive Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 119.8 to 132.1
Why not these models?Each has a cheaper-or-equal alternative that scores higher.
Kimi K3
Moonshot AI
Prefer Muse Spark 1.2. It costs less and leads by 0.8 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 108.7
- Cost
- $0.84/task
Audit
- Measured configuration
- Kimi K3 (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 102.7 to 112.8
- Gap to frontier
- 2.1 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
Fable 5
Anthropic
Prefer Opus 5. It costs less and leads by 3.6 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 122.3
- Cost
- $3.14/task
Audit
- Measured configuration
- Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 116.9 to 126.3
- Gap to frontier
- 3.6 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
- Note
- AA's harness used the documented Opus 4.8 fallback path; disclosed by AA.
Grok 4.6
xAI
Prefer Muse Spark 1.2. It costs less and leads by 6.1 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 103.4
- Cost
- $0.84/task
Audit
- Measured configuration
- Grok 4.6 (high)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 98.2 to 111.4
- Gap to frontier
- 7.4 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
Qwen3.8-Max
Alibaba
Better buy: Muse Spark 1.2. Muse Spark 1.2 is +7.8 Latent points and 65% cheaper.
- Latent Score
- 101.7
- Cost
- $1.13/task
Audit
- Measured configuration
- Qwen3.8 Max
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 95.8 to 106.2
- Gap to frontier
- 9.6 Latent points at this cost
Gemini 3.7 Flash
Google
Prefer Muse Spark 1.2. It leads by 11.1 Latent points at the same cost, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 98.4
- Cost
- $0.40/task
Audit
- Measured configuration
- Gemini 3.7 Flash (high)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 87.0 to 110.8
- Gap to frontier
- 11.1 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
GPT-5.6 Terra
OpenAI
Better buy: Muse Spark 1.2. Muse Spark 1.2 is +13.4 Latent points and 22% cheaper.
- Latent Score
- 96.1
- Cost
- $0.51/task
Audit
- Measured configuration
- GPT-5.6 Terra (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 91.7 to 101.8
- Gap to frontier
- 13.8 Latent points at this cost
- Note
- AA now measures Terra at max effort: the configuration the sheet pins (retrieved 2026-08-17).
DeepSeek V4 Flash
DeepSeek
Better buy: GPT-5.6 Luna. GPT-5.6 Luna is +15.0 Latent points and 57% cheaper.
- Latent Score
- 73.0
- Cost
- $0.11/task
Audit
- Measured configuration
- DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 67.3 to 81.2
- Gap to frontier
- 17.1 Latent points at this cost
- Note
- AA measures Flash 0731 at Max effort: the configuration the sheet pins.
Sonnet 5
Anthropic
Prefer DeepSeek V4 Pro. It costs less and leads by 2.3 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 89.8
- Cost
- $2.29/task
Audit
- Measured configuration
- Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Latent 95% interval
- 80.5 to 95.3
- Gap to frontier
- 35.6 Latent points at this cost
- Note
- AA Intelligence Index workload at standard $3/$15 list pricing (retrieved 2026-08-17).
Not ranked2 of 15 models
- GLM-5.3No Artificial Analysis cost measurement107.7score-cost/task
- Qwen3.8-27BNo Artificial Analysis cost measurement72score-cost/task
Every configuration · observed cost
Pick the exact reasoning effort
Most models ship several reasoning efforts at very different prices. Here every published configuration competes under the same rules: dominance first, ladder ordered by cost, and the same Latent ≥ 80 floor for the pick. Rungs below GPT-5.6 Luna (max) are cheaper still, but fall below the declared capability floor.
GPT-5.6 Luna (low)
OpenAI
- Latent Score
- 33.4
- Cost
- $0.009/task
Audit
- Measured configuration
- GPT-5.6 Luna (low)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
+$0.002/task · +12.0 Latent points · 1.2× the cost
GPT-5.6 Luna (medium)
OpenAI
- Latent Score
- 45.4
- Cost
- $0.011/task
Audit
- Measured configuration
- GPT-5.6 Luna (medium)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
+$0.011/task · +23.9 Latent points · 2× the cost
GPT-5.6 Luna (high)
OpenAI
- Latent Score
- 69.3
- Cost
- $0.022/task
Audit
- Measured configuration
- GPT-5.6 Luna (high)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
+$0.010/task · +10.6 Latent points · 1.5× the cost
GPT-5.6 Luna (xhigh)
OpenAI
- Latent Score
- 79.9
- Cost
- $0.032/task
Audit
- Measured configuration
- GPT-5.6 Luna (xhigh)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
+$0.015/task · +8.1 Latent points · 1.5× the cost
GPT-5.6 Luna (max)
OpenAI
Best bang for buck
- Latent Score
- 88.0
- Cost
- $0.047/task
Audit
- Measured configuration
- GPT-5.6 Luna (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Note
- AA now measures Luna at max effort: the configuration the sheet pins (retrieved 2026-08-17).
+$0.20/task · +4.1 Latent points · 5.3× the cost
DeepSeek V4 Pro (max)
DeepSeek
- Latent Score
- 92.1
- Cost
- $0.25/task
Audit
- Measured configuration
- DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Note
- AA measures the 0813 API snapshot at Max effort: the configuration the sheet pins.
+$0.15/task · +17.4 Latent points · 1.6× the cost
Muse Spark 1.2 (xhigh)
Meta
- Latent Score
- 109.5
- Cost
- $0.40/task
Audit
- Measured configuration
- Muse Spark 1.2 (xhigh)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
+$0.83/task · +9.4 Latent points · 3.1× the cost
Opus 5 (high)
Anthropic
- Latent Score
- 118.9
- Cost
- $1.23/task
Audit
- Measured configuration
- Claude Opus 5 (Adaptive Reasoning, High Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
+$0.57/task · +4.3 Latent points · 1.5× the cost
Opus 5 (xhigh)
Anthropic
- Latent Score
- 123.2
- Cost
- $1.80/task
Audit
- Measured configuration
- Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
+$0.54/task · +2.7 Latent points · 1.3× the cost
Opus 5 (max)
Anthropic
Maximum intelligence
- Latent Score
- 125.9
- Cost
- $2.34/task
Audit
- Measured configuration
- Claude Opus 5 (Adaptive Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
Why not these models?Each has a cheaper-or-equal alternative that scores higher.
Fable 5 (max)
Anthropic
Prefer Opus 5 (xhigh). It costs less and leads by 0.9 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 122.3
- Cost
- $3.14/task
Audit
- Measured configuration
- Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 3.6 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
- Note
- AA's harness used the documented Opus 4.8 fallback path; disclosed by AA.
Gemini 3.7 Flash (medium)
Google
Prefer DeepSeek V4 Pro (max). It costs less and leads by 3.6 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 88.5
- Cost
- $0.26/task
Audit
- Measured configuration
- Gemini 3.7 Flash (medium)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 5.1 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
Kimi K3 (max)
Moonshot AI
Prefer Muse Spark 1.2 (xhigh). It costs less and leads by 0.8 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 108.7
- Cost
- $0.84/task
Audit
- Measured configuration
- Kimi K3 (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 7.0 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
GPT-5.6 Sol (max)
OpenAI
Better buy: Opus 5 (high). Opus 5 (high) is +7.4 Latent points at the same cost.
- Latent Score
- 111.5
- Cost
- $1.23/task
Audit
- Measured configuration
- GPT-5.6 Sol (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 7.4 Latent points at this cost
Opus 5 (medium)
Anthropic
Prefer Muse Spark 1.2 (xhigh). It costs less and leads by 2.8 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 106.7
- Cost
- $0.72/task
Audit
- Measured configuration
- Claude Opus 5 (Adaptive Reasoning, Medium Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 7.7 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
Gemini 3.7 Flash (high)
Google
Prefer Muse Spark 1.2 (xhigh). It leads by 11.1 Latent points at the same cost, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 98.4
- Cost
- $0.40/task
Audit
- Measured configuration
- Gemini 3.7 Flash (high)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 11.1 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
GPT-5.6 Sol (xhigh)
OpenAI
Prefer Muse Spark 1.2 (xhigh). It costs less and leads by 5.6 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 103.9
- Cost
- $0.81/task
Audit
- Measured configuration
- GPT-5.6 Sol (xhigh)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 11.5 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
Gemini 3.7 Flash (low)
Google
Prefer GPT-5.6 Luna (xhigh). It costs less and leads by 0.4 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 79.5
- Cost
- $0.16/task
Audit
- Measured configuration
- Gemini 3.7 Flash (low)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 11.5 Latent points at this cost
Grok 4.6 (high)
xAI
Prefer Muse Spark 1.2 (xhigh). It costs less and leads by 6.1 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 103.4
- Cost
- $0.84/task
Audit
- Measured configuration
- Grok 4.6 (high)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 12.3 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
GPT-5.6 Sol (high)
OpenAI
Better buy: Muse Spark 1.2 (xhigh). Muse Spark 1.2 (xhigh) is +12.2 Latent points and 27% cheaper.
- Latent Score
- 97.3
- Cost
- $0.55/task
Audit
- Measured configuration
- GPT-5.6 Sol (high)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 14.9 Latent points at this cost
GPT-5.6 Terra (max)
OpenAI
Better buy: Muse Spark 1.2 (xhigh). Muse Spark 1.2 (xhigh) is +13.4 Latent points and 22% cheaper.
- Latent Score
- 96.1
- Cost
- $0.51/task
Audit
- Measured configuration
- GPT-5.6 Terra (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 15.4 Latent points at this cost
- Note
- AA now measures Terra at max effort: the configuration the sheet pins (retrieved 2026-08-17).
GPT-5.6 Sol (medium)
OpenAI
Prefer DeepSeek V4 Pro (max). It costs less and leads by 1.1 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 91.0
- Cost
- $0.37/task
Audit
- Measured configuration
- GPT-5.6 Sol (medium)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 15.6 Latent points at this cost
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
Qwen3.8-Max
Alibaba
Better buy: Muse Spark 1.2 (xhigh). Muse Spark 1.2 (xhigh) is +7.8 Latent points and 65% cheaper.
- Latent Score
- 101.7
- Cost
- $1.13/task
Audit
- Measured configuration
- Qwen3.8 Max
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 16.5 Latent points at this cost
DeepSeek V4 Flash (max)
DeepSeek
Better buy: GPT-5.6 Luna (xhigh). GPT-5.6 Luna (xhigh) is +6.9 Latent points and 71% cheaper.
- Latent Score
- 73.0
- Cost
- $0.11/task
Audit
- Measured configuration
- DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 17.1 Latent points at this cost
- Note
- AA measures Flash 0731 at Max effort: the configuration the sheet pins.
GPT-5.6 Terra (xhigh)
OpenAI
Better buy: GPT-5.6 Luna (max). GPT-5.6 Luna (max) is +5.7 Latent points and 85% cheaper.
- Latent Score
- 82.3
- Cost
- $0.31/task
Audit
- Measured configuration
- GPT-5.6 Terra (xhigh)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 17.8 Latent points at this cost
GPT-5.6 Sol (low)
OpenAI
Better buy: GPT-5.6 Luna (xhigh). GPT-5.6 Luna (xhigh) is +5.9 Latent points and 86% cheaper.
- Latent Score
- 74.0
- Cost
- $0.23/task
Audit
- Measured configuration
- GPT-5.6 Sol (low)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 17.9 Latent points at this cost
GPT-5.6 Terra (high)
OpenAI
Better buy: GPT-5.6 Luna (xhigh). GPT-5.6 Luna (xhigh) is +6.8 Latent points and 85% cheaper.
- Latent Score
- 73.1
- Cost
- $0.22/task
Audit
- Measured configuration
- GPT-5.6 Terra (high)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 18.7 Latent points at this cost
Kimi K3 (low)
Moonshot AI
Prefer GPT-5.6 Luna (high). It costs less and leads by 1.6 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 67.7
- Cost
- $0.24/task
Audit
- Measured configuration
- Kimi K3 (low)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 24.3 Latent points at this cost
Opus 5 (low)
Anthropic
Better buy: GPT-5.6 Luna (max). GPT-5.6 Luna (max) is +4.6 Latent points and 89% cheaper.
- Latent Score
- 83.4
- Cost
- $0.43/task
Audit
- Measured configuration
- Claude Opus 5 (Adaptive Reasoning, Low Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 26.7 Latent points at this cost
GPT-5.6 Terra (medium)
OpenAI
Better buy: GPT-5.6 Luna (high). GPT-5.6 Luna (high) is +6.5 Latent points and 82% cheaper.
- Latent Score
- 62.8
- Cost
- $0.12/task
Audit
- Measured configuration
- GPT-5.6 Terra (medium)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 27.5 Latent points at this cost
GPT-5.6 Luna (non-reasoning)
OpenAI
Better buy: GPT-5.6 Luna (low). GPT-5.6 Luna (low) is +13.7 Latent points and 25% cheaper.
- Latent Score
- 19.7
- Cost
- $0.012/task
Audit
- Measured configuration
- GPT-5.6 Luna (Non-reasoning)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 28.7 Latent points at this cost
Sonnet 5 (max)
Anthropic
Prefer DeepSeek V4 Pro (max). It costs less and leads by 2.3 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 89.8
- Cost
- $2.29/task
Audit
- Measured configuration
- Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 35.9 Latent points at this cost
- Note
- AA Intelligence Index workload at standard $3/$15 list pricing (retrieved 2026-08-17).
GPT-5.6 Terra (low)
OpenAI
Better buy: GPT-5.6 Luna (high). GPT-5.6 Luna (high) is +21.8 Latent points and 77% cheaper.
- Latent Score
- 47.5
- Cost
- $0.094/task
Audit
- Measured configuration
- GPT-5.6 Terra (low)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 42.2 Latent points at this cost
GPT-5.6 Sol (non-reasoning)
OpenAI
Better buy: GPT-5.6 Luna (high). GPT-5.6 Luna (high) is +21.0 Latent points and 91% cheaper.
- Latent Score
- 48.3
- Cost
- $0.24/task
Audit
- Measured configuration
- GPT-5.6 Sol (Non-reasoning)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 43.7 Latent points at this cost
GPT-5.6 Terra (non-reasoning)
OpenAI
Better buy: GPT-5.6 Luna (low). GPT-5.6 Luna (low) is +1.4 Latent points and 91% cheaper.
- Latent Score
- 32.0
- Cost
- $0.10/task
Audit
- Measured configuration
- GPT-5.6 Terra (Non-reasoning)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Gap to frontier
- 57.9 Latent points at this cost
Not ranked2 of 37 models
- GLM-5.3No Artificial Analysis cost measurement107.7score-cost/task
- Qwen3.8-27BNo Artificial Analysis cost measurement72score-cost/task
37 published configurations across 15 models. Observed cost per task by Artificial Analysis (Intelligence Index v4.1.1 workload, retrieved 2026-08-17). Models without a published effort knob appear once.
Best value, ranked · observed cost
Every configuration, best value first
No scalar value score decides this order. Configurations clearing the Latent ≥ 80 floor rank ahead of those below it; within each group, by distance from the cost-capability frontier (0 = no better trade-off exists at its price); ties go to the cheaper configuration.
| # | Configuration | Latent Score | Cost/task | Gap to frontier | Verdict |
|---|---|---|---|---|---|
| 1 | GPT-5.6 Luna (max)OpenAI | 88.0 | $0.047 | 0 | Best value: the cheapest configuration clearing the capability floor |
| 2 | DeepSeek V4 Pro (max)DeepSeek | 92.1 | $0.25 | 0 | Step-up option: unbeatable at its price |
| 3 | Muse Spark 1.2 (xhigh)Meta | 109.5 | $0.40 | 0 | Step-up option: unbeatable at its price |
| 4 | Opus 5 (high)Anthropic | 118.9 | $1.23 | 0 | Step-up option: unbeatable at its price |
| 5 | Opus 5 (xhigh)Anthropic | 123.2 | $1.80 | 0 | Step-up option: unbeatable at its price |
| 6 | Opus 5 (max)Anthropic | 125.9 | $2.34 | 0 | Maximum intelligence: the strongest brain on this board. Worth the multiple only for the most complex reasoning work the benchmarks reward. |
| 7 | Fable 5 (max)Anthropic | 122.3 | $3.14 | −3.6 | Prefer Opus 5 (xhigh); pay up only for the hardest reasoning (gap within uncertainty) |
| 8 | Gemini 3.7 Flash (medium)Google | 88.5 | $0.26 | −5.1 | Prefer DeepSeek V4 Pro (max); pay up only for the hardest reasoning (gap within uncertainty) |
| 9 | Kimi K3 (max)Moonshot AI | 108.7 | $0.84 | −7.0 | Prefer Muse Spark 1.2 (xhigh); pay up only for the hardest reasoning (gap within uncertainty) |
| 10 | GPT-5.6 Sol (max)OpenAI | 111.5 | $1.23 | −7.4 | Better buy: Opus 5 (high). +7.4 points at the same cost |
| 11 | Opus 5 (medium)Anthropic | 106.7 | $0.72 | −7.7 | Prefer Muse Spark 1.2 (xhigh); pay up only for the hardest reasoning (gap within uncertainty) |
| 12 | Gemini 3.7 Flash (high)Google | 98.4 | $0.40 | −11.1 | Prefer Muse Spark 1.2 (xhigh); pay up only for the hardest reasoning (gap within uncertainty) |
| 13 | GPT-5.6 Sol (xhigh)OpenAI | 103.9 | $0.81 | −11.5 | Prefer Muse Spark 1.2 (xhigh); pay up only for the hardest reasoning (gap within uncertainty) |
| 14 | Grok 4.6 (high)xAI | 103.4 | $0.84 | −12.3 | Prefer Muse Spark 1.2 (xhigh); pay up only for the hardest reasoning (gap within uncertainty) |
| 15 | GPT-5.6 Sol (high)OpenAI | 97.3 | $0.55 | −14.9 | Better buy: Muse Spark 1.2 (xhigh). +12.2 points, 27% cheaper |
| 16 | GPT-5.6 Terra (max)OpenAI | 96.1 | $0.51 | −15.4 | Better buy: Muse Spark 1.2 (xhigh). +13.4 points, 22% cheaper |
| 17 | GPT-5.6 Sol (medium)OpenAI | 91.0 | $0.37 | −15.6 | Prefer DeepSeek V4 Pro (max); pay up only for the hardest reasoning (gap within uncertainty) |
| 18 | Qwen3.8-MaxAlibaba | 101.7 | $1.13 | −16.5 | Better buy: Muse Spark 1.2 (xhigh). +7.8 points, 65% cheaper |
| 19 | GPT-5.6 Terra (xhigh)OpenAI | 82.3 | $0.31 | −17.8 | Better buy: GPT-5.6 Luna (max). +5.7 points, 85% cheaper |
| 20 | Opus 5 (low)Anthropic | 83.4 | $0.43 | −26.7 | Better buy: GPT-5.6 Luna (max). +4.6 points, 89% cheaper |
| 21 | Sonnet 5 (max)Anthropic | 89.8 | $2.29 | −35.9 | Prefer DeepSeek V4 Pro (max); pay up only for the hardest reasoning (gap within uncertainty) |
| 22 | GPT-5.6 Luna (low)OpenAI | 33.4 | $0.009 | 0 | Unbeatable at its price, but below the Latent ≥ 80 floor |
| 23 | GPT-5.6 Luna (medium)OpenAI | 45.4 | $0.011 | 0 | Unbeatable at its price, but below the Latent ≥ 80 floor |
| 24 | GPT-5.6 Luna (high)OpenAI | 69.3 | $0.022 | 0 | Unbeatable at its price, but below the Latent ≥ 80 floor |
| 25 | GPT-5.6 Luna (xhigh)OpenAI | 79.9 | $0.032 | 0 | Unbeatable at its price, but below the Latent ≥ 80 floor |
| 26 | Gemini 3.7 Flash (low)Google | 79.5 | $0.16 | −11.5 | Prefer GPT-5.6 Luna (xhigh); pay up only for the hardest reasoning (gap within uncertainty) |
| 27 | DeepSeek V4 Flash (max)DeepSeek | 73.0 | $0.11 | −17.1 | Better buy: GPT-5.6 Luna (xhigh). +6.9 points, 71% cheaper |
| 28 | GPT-5.6 Sol (low)OpenAI | 74.0 | $0.23 | −17.9 | Better buy: GPT-5.6 Luna (xhigh). +5.9 points, 86% cheaper |
| 29 | GPT-5.6 Terra (high)OpenAI | 73.1 | $0.22 | −18.7 | Better buy: GPT-5.6 Luna (xhigh). +6.8 points, 85% cheaper |
| 30 | Kimi K3 (low)Moonshot AI | 67.7 | $0.24 | −24.3 | Prefer GPT-5.6 Luna (high); pay up only for the hardest reasoning (gap within uncertainty) |
| 31 | GPT-5.6 Terra (medium)OpenAI | 62.8 | $0.12 | −27.5 | Better buy: GPT-5.6 Luna (high). +6.5 points, 82% cheaper |
| 32 | GPT-5.6 Luna (non-reasoning)OpenAI | 19.7 | $0.012 | −28.7 | Better buy: GPT-5.6 Luna (low). +13.7 points, 25% cheaper |
| 33 | GPT-5.6 Terra (low)OpenAI | 47.5 | $0.094 | −42.2 | Better buy: GPT-5.6 Luna (high). +21.8 points, 77% cheaper |
| 34 | GPT-5.6 Sol (non-reasoning)OpenAI | 48.3 | $0.24 | −43.7 | Better buy: GPT-5.6 Luna (high). +21.0 points, 91% cheaper |
| 35 | GPT-5.6 Terra (non-reasoning)OpenAI | 32.0 | $0.10 | −57.9 | Better buy: GPT-5.6 Luna (low). +1.4 points, 91% cheaper |
A low rank is not a ban: the most capable configurations are unbeatable at what they do. The ranking answers "is there a better way to spend your money?"; the verdict says when there isn't. Provisional models are never ranked.
Secondary lens · listed price, fixed basket
Compare by listed token prices
Same purchasing rules, different definition of cost: 1,000,000 input + 250,000 output tokens at listed first-party rates. Reasoning-token consumption is invisible here. That is what the observed-cost view above adds.
DeepSeek V4 Flash
DeepSeek
- Latent Score
- 73.0
- Cost
- $0.21 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $0.14/M in · $0.28/M out
- Price updated
- 2026-08-17
- Latent 95% interval
- 67.3 to 81.2
- Note
- Cache-miss list price for deepseek-v4-flash (0731 checkpoint).
+$0.29 basket · +15.0 Latent points · 2.4× the cost
GPT-5.6 Luna
OpenAI
Best bang for buck
- Latent Score
- 88.0
- Cost
- $0.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $0.2/M in · $1.2/M out
- Price updated
- 2026-08-17
- Latent 95% interval
- 83.2 to 95.0
- Note
- Standard API pricing.
+$1.19 basket · +10.4 Latent points · 3.4× the cost
Gemini 3.7 Flash
Google
- Latent Score
- 98.4
- Cost
- $1.69 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $0.75/M in · $3.75/M out
- Price updated
- 2026-08-16
- Latent 95% interval
- 87.0 to 110.8
- Note
- Introductory rate through 2026-12-31; standard rate is $1.50/$7.50 from 2027-01-01.
- Disclosed rate scenario
- Standard rate from 2027-01-01 (introductory rate expires): basket $3.38
+$0.63 basket · +11.1 Latent points · 1.4× the cost
Muse Spark 1.2
Meta
- Latent Score
- 109.5
- Cost
- $2.31 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $1.25/M in · $4.25/M out
- Price updated
- 2026-08-16
- Latent 95% interval
- 98.7 to 115.7
- Note
- Standard tier; contributor tier is excluded (training-data exchange).
+$8.94 basket · +16.4 Latent points · 4.9× the cost
Opus 5
Anthropic
Maximum intelligence
- Latent Score
- 125.9
- Cost
- $11.25 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $5/M in · $25/M out
- Price updated
- 2026-08-15
- Latent 95% interval
- 119.8 to 132.1
- Note
- Standard uncached API pricing.
Why not these models?Each has a cheaper-or-equal alternative that scores higher.
Fable 5
Anthropic
Prefer Opus 5. It costs less and leads by 3.6 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 122.3
- Cost
- $22.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $10/M in · $50/M out
- Price updated
- 2026-08-15
- Latent 95% interval
- 116.9 to 126.3
- Gap to frontier
- 3.6 Latent points at this price
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
- Note
- Standard uncached API pricing.
Grok 4.6
xAI
Prefer Muse Spark 1.2. It costs less and leads by 6.1 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 103.4
- Cost
- $3.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $6/M out
- Price updated
- 2026-08-15
- Latent 95% interval
- 98.2 to 111.4
- Gap to frontier
- 10.4 Latent points at this price
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
- Note
- Base tier below 200k prompt tokens.
Kimi K3
Moonshot AI
Prefer Muse Spark 1.2. It costs less and leads by 0.8 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 108.7
- Cost
- $6.75 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $3/M in · $15/M out
- Price updated
- 2026-08-15
- Latent 95% interval
- 102.7 to 112.8
- Gap to frontier
- 11.9 Latent points at this price
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
- Note
- Hosted API cache-miss pricing.
Qwen3.8-Max
Alibaba
Better buy: Muse Spark 1.2. Muse Spark 1.2 is +7.8 Latent points and 34% cheaper.
- Latent Score
- 101.7
- Cost
- $3.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $6/M out
- Price updated
- 2026-08-16
- Latent 95% interval
- 95.8 to 106.2
- Gap to frontier
- 12.1 Latent points at this price
- Note
- International (Singapore) Model Studio list pricing.
GPT-5.6 Sol
OpenAI
Better buy: Opus 5. Opus 5 is +14.4 Latent points and 10% cheaper.
- Latent Score
- 111.5
- Cost
- $12.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $5/M in · $30/M out
- Price updated
- 2026-08-15
- Latent 95% interval
- 106.1 to 116.6
- Gap to frontier
- 14.4 Latent points at this price
- Note
- Standard tier below the long-context threshold.
DeepSeek V4 Pro
DeepSeek
Prefer Gemini 3.7 Flash. It costs less and leads by 6.3 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 92.1
- Cost
- $2.31 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $1.32/M in · $3.96/M out
- Price updated
- 2026-08-16
- Latent 95% interval
- 85.8 to 100.8
- Gap to frontier
- 17.4 Latent points at this price
- Robust frontier
- Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
- Note
- Peak cache-miss rate effective 2026-08-16; off-peak is half of peak.
- Disclosed rate scenario
- Off-peak rate (published; half of peak): basket $1.16
GPT-5.6 Terra
OpenAI
Prefer Gemini 3.7 Flash. It costs less and leads by 2.3 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.
- Latent Score
- 96.1
- Cost
- $5.00 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $12/M out
- Price updated
- 2026-08-17
- Latent 95% interval
- 91.7 to 101.8
- Gap to frontier
- 21.4 Latent points at this price
- Note
- Standard tier below the long-context threshold.
Sonnet 5
Anthropic
Better buy: Gemini 3.7 Flash. Gemini 3.7 Flash is +8.6 Latent points and 62% cheaper.
- Latent Score
- 89.8
- Cost
- $4.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $10/M out
- Price updated
- 2026-08-23
- Latent 95% interval
- 80.5 to 95.3
- Gap to frontier
- 26.6 Latent points at this price
- Note
- Standard API pricing.
Not ranked2 of 15 models
- GLM-5.3No public per-token price107.7score-basket
- Qwen3.8-27BNo public per-token price72score-basket
le-2026-v2.0 · lv-2026-v3.0
Rule-based purchasing criteria, two cost lenses
Observed cost, strictly matched
The headline view uses Artificial Analysis's measured cost to complete one Intelligence Index v4.1.1 task (input, cache, reasoning, and answer tokens included). The model board accepts only the exact configuration the sheet pins; a mismatch or missing measurement means unranked, never approximated. The configuration board applies the same rules to every published reasoning effort of the ranked models. The frontier chart additionally plots estimated effort variants for models Artificial Analysis measures only once (Fable 5, Grok 4.6, Qwen3.8-Max, Muse Spark 1.2): scores derive from CursorBench, the one public benchmark that tests every reasoning effort, calibrated on the measured effort ladders; costs derive from the measured cost ladders. Each is marked as an estimate and never ranked.
Dominance before everything
A model that another ranked model beats on both dimensions (cheaper or equal in cost, at least as capable, and strictly better on one) is never recommended ahead of its dominator. Point estimates decide placement; the bootstrap decides wording. A dominance claim is presented as decisive only when the dominator's capability advantage is positive in ≥90% of Latent bootstrap replicates. Cost is deterministic, so only capability needs the test.
A declared capability floor
The Value Pick is the cheapest non-dominated model clearing Latent ≥ 80, so an extremely cheap but weak model can never headline on price alone. Declared minimum useful-capability threshold: 1.33 SD below the 100-point cohort anchor on the 100±15 scale. A versioned policy choice, re-declared whenever the scale re-anchors; not discovered from the data.
A ladder, not a value race
Models above the pick are not "worse value" in some scalar sense. They are step-up options, ordered by cost, each stating exactly what the extra spend buys in Latent points. The scalar Efficiency and Value scores published under the previous board versions remain in the released snapshots for reproducibility, but no displayed ordering uses them.
Observed cost per task measurements by Artificial Analysis (Intelligence Index v4.1.1 workload, retrieved 2026-08-17), used as cost evidence only. Latency, throughput, reliability, and human preference are not silently folded into any headline.