Latent Value & Efficiency

Which model gives you the most intelligence for the money?

A rule-based buying guide: cheaper and smarter always wins, and the headline pick is the cheapest genuinely capable model. Cost is what a model actually spent per benchmark task, measured independently by Artificial Analysis — reasoning tokens included, and price never feeds back into the intelligence estimate.

Best bang for your buck

GPT-5.6 Luna

No cheaper model comes close; every smarter model costs multiples more.

Latent Score
88.0
Observed cost / task
$0.047
vs DeepSeek V4 Flash
57% cheaper

+15.0 Latent points

Capability × observed cost

The Latent Frontier

Up and to the left is better. Filled red points are the buying ladder: no cheaper option beats them. Every published reasoning configuration is plotted.

Latent Frontier: score against observed cost per taskEvery published reasoning configuration plotted as Latent Score against observed cost per benchmark task, as measured by Artificial Analysis. Buying-ladder configurations are highlighted; hollow points are provisional; dashed points are estimates derived from public effort benchmarks and are never ranked.36547290108126$0.01$0.02$0.05$0.10$0.15$0.25$0.40$0.60$1$1.5$2.5$4Opus 5 (max)Opus 5 (xhigh)Opus 5 (high)Opus 5 (medium)Opus 5 (low)Fable 5 (max)5.6 Sol (max)5.6 Sol (xhigh)5.6 Sol (high)5.6 Sol (medium)5.6 Sol (low)Muse Spark 1.2 (xhigh)Kimi K3 (max)Kimi K3 (low)Grok 4.6 (high)Qwen3.8-MaxGemini 3.7 Flash (high)Gemini 3.7 Flash (medium)Gemini 3.7 Flash (low)5.6 Terra (max)5.6 Terra (xhigh)5.6 Terra (high)5.6 Terra (medium)5.6 Terra (low)DeepSeek V4 Pro (max)Sonnet 5 (max)5.6 Luna (max)5.6 Luna (xhigh)5.6 Luna (high)5.6 Luna (medium)5.6 Luna (low)DeepSeek V4 Flash (max)Fable 5 (xhigh) est.Fable 5 (high) est.Fable 5 (medium) est.Fable 5 (low) est.Grok 4.6 (xhigh) est.Grok 4.6 (medium) est.Grok 4.6 (low) est.Muse Spark 1.2 (high) est.Muse Spark 1.2 (medium) est.Muse Spark 1.2 (low) est.Qwen3.8-Max (medium) est.Qwen3.8-Max (low) est.GLM-5.3 est.Observed cost per Intelligence Index task · log scaleThe Latent Score
  • Buying ladder
  • Dominated, within uncertainty of the ladder
  • Dominated
  • Estimated (est.): derived from public effort benchmarks, never ranked

Qwen3.8-27B has no Artificial Analysis cost measurement and cannot be plotted. GLM-5.3 has no Artificial Analysis cost measurement; its point uses an estimated cost transferred from its predecessor's measured rate at price parity. Cost data: Artificial Analysis Intelligence Index v4.1.1, retrieved 2026-08-17.

Observed cost · Artificial Analysis

Spend more for more intelligence

No scalar value score. A model dominated by a cheaper-and-smarter model is never recommended ahead of it; non-dominated models form a buying ladder, cheapest first; the Value Pick is the cheapest ladder model clearing the declared Latent ≥ 80 floor.

  1. GPT-5.6 Luna

    OpenAI

    Best bang for buck

    Latent Score
    88.0
    Cost
    $0.047/task
    Audit
    Measured configuration
    GPT-5.6 Luna (max)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    83.2 to 95.0
    Note
    AA now measures Luna at max effort: the configuration the sheet pins (retrieved 2026-08-17).
    Open cost source →View The Latent Score →
  2. +$0.20/task · +4.1 Latent points · 5.3× the cost

    DeepSeek V4 Pro

    DeepSeek

    Latent Score
    92.1
    Cost
    $0.25/task
    Audit
    Measured configuration
    DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    85.8 to 100.8
    Note
    AA measures the 0813 API snapshot at Max effort: the configuration the sheet pins.
    Open cost source →View The Latent Score →
  3. +$0.15/task · +17.4 Latent points · 1.6× the cost

    Muse Spark 1.2

    Meta

    Latent Score
    109.5
    Cost
    $0.40/task
    Audit
    Measured configuration
    Muse Spark 1.2 (xhigh)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    98.7 to 115.7
    Open cost source →View The Latent Score →
  4. +$0.83/task · +2.0 Latent points · 3.1× the cost

    GPT-5.6 Sol

    OpenAI

    Latent Score
    111.5
    Cost
    $1.23/task
    Audit
    Measured configuration
    GPT-5.6 Sol (max)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    106.1 to 116.6
    Open cost source →View The Latent Score →
  5. +$1.11/task · +14.4 Latent points · 1.9× the cost

    Opus 5

    Anthropic

    Maximum intelligence

    Latent Score
    125.9
    Cost
    $2.34/task
    Audit
    Measured configuration
    Claude Opus 5 (Adaptive Reasoning, Max Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    119.8 to 132.1
    Open cost source →View The Latent Score →

Why not these models?Each has a cheaper-or-equal alternative that scores higher.

  • Kimi K3

    Moonshot AI

    Prefer Muse Spark 1.2. It costs less and leads by 0.8 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    108.7
    Cost
    $0.84/task
    Audit
    Measured configuration
    Kimi K3 (max)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    102.7 to 112.8
    Gap to frontier
    2.1 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
    Open cost source →View The Latent Score →
  • Fable 5

    Anthropic

    Prefer Opus 5. It costs less and leads by 3.6 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    122.3
    Cost
    $3.14/task
    Audit
    Measured configuration
    Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    116.9 to 126.3
    Gap to frontier
    3.6 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
    Note
    AA's harness used the documented Opus 4.8 fallback path; disclosed by AA.
    Open cost source →View The Latent Score →
  • Grok 4.6

    xAI

    Prefer Muse Spark 1.2. It costs less and leads by 6.1 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    103.4
    Cost
    $0.84/task
    Audit
    Measured configuration
    Grok 4.6 (high)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    98.2 to 111.4
    Gap to frontier
    7.4 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
    Open cost source →View The Latent Score →
  • Qwen3.8-Max

    Alibaba

    Better buy: Muse Spark 1.2. Muse Spark 1.2 is +7.8 Latent points and 65% cheaper.

    Latent Score
    101.7
    Cost
    $1.13/task
    Audit
    Measured configuration
    Qwen3.8 Max
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    95.8 to 106.2
    Gap to frontier
    9.6 Latent points at this cost
    Open cost source →View The Latent Score →
  • Gemini 3.7 Flash

    Google

    Prefer Muse Spark 1.2. It leads by 11.1 Latent points at the same cost, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    98.4
    Cost
    $0.40/task
    Audit
    Measured configuration
    Gemini 3.7 Flash (high)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    87.0 to 110.8
    Gap to frontier
    11.1 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
    Open cost source →View The Latent Score →
  • GPT-5.6 Terra

    OpenAI

    Better buy: Muse Spark 1.2. Muse Spark 1.2 is +13.4 Latent points and 22% cheaper.

    Latent Score
    96.1
    Cost
    $0.51/task
    Audit
    Measured configuration
    GPT-5.6 Terra (max)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    91.7 to 101.8
    Gap to frontier
    13.8 Latent points at this cost
    Note
    AA now measures Terra at max effort: the configuration the sheet pins (retrieved 2026-08-17).
    Open cost source →View The Latent Score →
  • DeepSeek V4 Flash

    DeepSeek

    Better buy: GPT-5.6 Luna. GPT-5.6 Luna is +15.0 Latent points and 57% cheaper.

    Latent Score
    73.0
    Cost
    $0.11/task
    Audit
    Measured configuration
    DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    67.3 to 81.2
    Gap to frontier
    17.1 Latent points at this cost
    Note
    AA measures Flash 0731 at Max effort: the configuration the sheet pins.
    Open cost source →View The Latent Score →
  • Sonnet 5

    Anthropic

    Prefer DeepSeek V4 Pro. It costs less and leads by 2.3 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    89.8
    Cost
    $2.29/task
    Audit
    Measured configuration
    Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Latent 95% interval
    80.5 to 95.3
    Gap to frontier
    35.6 Latent points at this cost
    Note
    AA Intelligence Index workload at standard $3/$15 list pricing (retrieved 2026-08-17).
    Open cost source →View The Latent Score →

Not ranked2 of 15 models

  • GLM-5.3No Artificial Analysis cost measurement107.7score-cost/task
  • Qwen3.8-27BNo Artificial Analysis cost measurement72score-cost/task

Every configuration · observed cost

Pick the exact reasoning effort

Most models ship several reasoning efforts at very different prices. Here every published configuration competes under the same rules: dominance first, ladder ordered by cost, and the same Latent ≥ 80 floor for the pick. Rungs below GPT-5.6 Luna (max) are cheaper still, but fall below the declared capability floor.

  1. GPT-5.6 Luna (low)

    OpenAI

    Latent Score
    33.4
    Cost
    $0.009/task
    Audit
    Measured configuration
    GPT-5.6 Luna (low)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Open cost source →
  2. +$0.002/task · +12.0 Latent points · 1.2× the cost

    GPT-5.6 Luna (medium)

    OpenAI

    Latent Score
    45.4
    Cost
    $0.011/task
    Audit
    Measured configuration
    GPT-5.6 Luna (medium)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Open cost source →
  3. +$0.011/task · +23.9 Latent points · 2× the cost

    GPT-5.6 Luna (high)

    OpenAI

    Latent Score
    69.3
    Cost
    $0.022/task
    Audit
    Measured configuration
    GPT-5.6 Luna (high)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Open cost source →
  4. +$0.010/task · +10.6 Latent points · 1.5× the cost

    GPT-5.6 Luna (xhigh)

    OpenAI

    Latent Score
    79.9
    Cost
    $0.032/task
    Audit
    Measured configuration
    GPT-5.6 Luna (xhigh)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Open cost source →
  5. +$0.015/task · +8.1 Latent points · 1.5× the cost

    GPT-5.6 Luna (max)

    OpenAI

    Best bang for buck

    Latent Score
    88.0
    Cost
    $0.047/task
    Audit
    Measured configuration
    GPT-5.6 Luna (max)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Note
    AA now measures Luna at max effort: the configuration the sheet pins (retrieved 2026-08-17).
    Open cost source →View The Latent Score →
  6. +$0.20/task · +4.1 Latent points · 5.3× the cost

    DeepSeek V4 Pro (max)

    DeepSeek

    Latent Score
    92.1
    Cost
    $0.25/task
    Audit
    Measured configuration
    DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Note
    AA measures the 0813 API snapshot at Max effort: the configuration the sheet pins.
    Open cost source →View The Latent Score →
  7. +$0.15/task · +17.4 Latent points · 1.6× the cost

    Muse Spark 1.2 (xhigh)

    Meta

    Latent Score
    109.5
    Cost
    $0.40/task
    Audit
    Measured configuration
    Muse Spark 1.2 (xhigh)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Open cost source →View The Latent Score →
  8. +$0.83/task · +9.4 Latent points · 3.1× the cost

    Opus 5 (high)

    Anthropic

    Latent Score
    118.9
    Cost
    $1.23/task
    Audit
    Measured configuration
    Claude Opus 5 (Adaptive Reasoning, High Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Open cost source →
  9. +$0.57/task · +4.3 Latent points · 1.5× the cost

    Opus 5 (xhigh)

    Anthropic

    Latent Score
    123.2
    Cost
    $1.80/task
    Audit
    Measured configuration
    Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Open cost source →
  10. +$0.54/task · +2.7 Latent points · 1.3× the cost

    Opus 5 (max)

    Anthropic

    Maximum intelligence

    Latent Score
    125.9
    Cost
    $2.34/task
    Audit
    Measured configuration
    Claude Opus 5 (Adaptive Reasoning, Max Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Open cost source →View The Latent Score →

Why not these models?Each has a cheaper-or-equal alternative that scores higher.

  • Fable 5 (max)

    Anthropic

    Prefer Opus 5 (xhigh). It costs less and leads by 0.9 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    122.3
    Cost
    $3.14/task
    Audit
    Measured configuration
    Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    3.6 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
    Note
    AA's harness used the documented Opus 4.8 fallback path; disclosed by AA.
    Open cost source →View The Latent Score →
  • Gemini 3.7 Flash (medium)

    Google

    Prefer DeepSeek V4 Pro (max). It costs less and leads by 3.6 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    88.5
    Cost
    $0.26/task
    Audit
    Measured configuration
    Gemini 3.7 Flash (medium)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    5.1 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
    Open cost source →
  • Kimi K3 (max)

    Moonshot AI

    Prefer Muse Spark 1.2 (xhigh). It costs less and leads by 0.8 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    108.7
    Cost
    $0.84/task
    Audit
    Measured configuration
    Kimi K3 (max)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    7.0 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
    Open cost source →View The Latent Score →
  • GPT-5.6 Sol (max)

    OpenAI

    Better buy: Opus 5 (high). Opus 5 (high) is +7.4 Latent points at the same cost.

    Latent Score
    111.5
    Cost
    $1.23/task
    Audit
    Measured configuration
    GPT-5.6 Sol (max)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    7.4 Latent points at this cost
    Open cost source →View The Latent Score →
  • Opus 5 (medium)

    Anthropic

    Prefer Muse Spark 1.2 (xhigh). It costs less and leads by 2.8 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    106.7
    Cost
    $0.72/task
    Audit
    Measured configuration
    Claude Opus 5 (Adaptive Reasoning, Medium Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    7.7 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
    Open cost source →
  • Gemini 3.7 Flash (high)

    Google

    Prefer Muse Spark 1.2 (xhigh). It leads by 11.1 Latent points at the same cost, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    98.4
    Cost
    $0.40/task
    Audit
    Measured configuration
    Gemini 3.7 Flash (high)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    11.1 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
    Open cost source →View The Latent Score →
  • GPT-5.6 Sol (xhigh)

    OpenAI

    Prefer Muse Spark 1.2 (xhigh). It costs less and leads by 5.6 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    103.9
    Cost
    $0.81/task
    Audit
    Measured configuration
    GPT-5.6 Sol (xhigh)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    11.5 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
    Open cost source →
  • Gemini 3.7 Flash (low)

    Google

    Prefer GPT-5.6 Luna (xhigh). It costs less and leads by 0.4 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    79.5
    Cost
    $0.16/task
    Audit
    Measured configuration
    Gemini 3.7 Flash (low)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    11.5 Latent points at this cost
    Open cost source →
  • Grok 4.6 (high)

    xAI

    Prefer Muse Spark 1.2 (xhigh). It costs less and leads by 6.1 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    103.4
    Cost
    $0.84/task
    Audit
    Measured configuration
    Grok 4.6 (high)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    12.3 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
    Open cost source →View The Latent Score →
  • GPT-5.6 Sol (high)

    OpenAI

    Better buy: Muse Spark 1.2 (xhigh). Muse Spark 1.2 (xhigh) is +12.2 Latent points and 27% cheaper.

    Latent Score
    97.3
    Cost
    $0.55/task
    Audit
    Measured configuration
    GPT-5.6 Sol (high)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    14.9 Latent points at this cost
    Open cost source →
  • GPT-5.6 Terra (max)

    OpenAI

    Better buy: Muse Spark 1.2 (xhigh). Muse Spark 1.2 (xhigh) is +13.4 Latent points and 22% cheaper.

    Latent Score
    96.1
    Cost
    $0.51/task
    Audit
    Measured configuration
    GPT-5.6 Terra (max)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    15.4 Latent points at this cost
    Note
    AA now measures Terra at max effort: the configuration the sheet pins (retrieved 2026-08-17).
    Open cost source →View The Latent Score →
  • GPT-5.6 Sol (medium)

    OpenAI

    Prefer DeepSeek V4 Pro (max). It costs less and leads by 1.1 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    91.0
    Cost
    $0.37/task
    Audit
    Measured configuration
    GPT-5.6 Sol (medium)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    15.6 Latent points at this cost
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal configuration is more capable at ≥90% bootstrap probability.
    Open cost source →
  • Qwen3.8-Max

    Alibaba

    Better buy: Muse Spark 1.2 (xhigh). Muse Spark 1.2 (xhigh) is +7.8 Latent points and 65% cheaper.

    Latent Score
    101.7
    Cost
    $1.13/task
    Audit
    Measured configuration
    Qwen3.8 Max
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    16.5 Latent points at this cost
    Open cost source →View The Latent Score →
  • DeepSeek V4 Flash (max)

    DeepSeek

    Better buy: GPT-5.6 Luna (xhigh). GPT-5.6 Luna (xhigh) is +6.9 Latent points and 71% cheaper.

    Latent Score
    73.0
    Cost
    $0.11/task
    Audit
    Measured configuration
    DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    17.1 Latent points at this cost
    Note
    AA measures Flash 0731 at Max effort: the configuration the sheet pins.
    Open cost source →View The Latent Score →
  • GPT-5.6 Terra (xhigh)

    OpenAI

    Better buy: GPT-5.6 Luna (max). GPT-5.6 Luna (max) is +5.7 Latent points and 85% cheaper.

    Latent Score
    82.3
    Cost
    $0.31/task
    Audit
    Measured configuration
    GPT-5.6 Terra (xhigh)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    17.8 Latent points at this cost
    Open cost source →
  • GPT-5.6 Sol (low)

    OpenAI

    Better buy: GPT-5.6 Luna (xhigh). GPT-5.6 Luna (xhigh) is +5.9 Latent points and 86% cheaper.

    Latent Score
    74.0
    Cost
    $0.23/task
    Audit
    Measured configuration
    GPT-5.6 Sol (low)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    17.9 Latent points at this cost
    Open cost source →
  • GPT-5.6 Terra (high)

    OpenAI

    Better buy: GPT-5.6 Luna (xhigh). GPT-5.6 Luna (xhigh) is +6.8 Latent points and 85% cheaper.

    Latent Score
    73.1
    Cost
    $0.22/task
    Audit
    Measured configuration
    GPT-5.6 Terra (high)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    18.7 Latent points at this cost
    Open cost source →
  • Kimi K3 (low)

    Moonshot AI

    Prefer GPT-5.6 Luna (high). It costs less and leads by 1.6 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    67.7
    Cost
    $0.24/task
    Audit
    Measured configuration
    Kimi K3 (low)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    24.3 Latent points at this cost
    Open cost source →
  • Opus 5 (low)

    Anthropic

    Better buy: GPT-5.6 Luna (max). GPT-5.6 Luna (max) is +4.6 Latent points and 89% cheaper.

    Latent Score
    83.4
    Cost
    $0.43/task
    Audit
    Measured configuration
    Claude Opus 5 (Adaptive Reasoning, Low Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    26.7 Latent points at this cost
    Open cost source →
  • GPT-5.6 Terra (medium)

    OpenAI

    Better buy: GPT-5.6 Luna (high). GPT-5.6 Luna (high) is +6.5 Latent points and 82% cheaper.

    Latent Score
    62.8
    Cost
    $0.12/task
    Audit
    Measured configuration
    GPT-5.6 Terra (medium)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    27.5 Latent points at this cost
    Open cost source →
  • GPT-5.6 Luna (non-reasoning)

    OpenAI

    Better buy: GPT-5.6 Luna (low). GPT-5.6 Luna (low) is +13.7 Latent points and 25% cheaper.

    Latent Score
    19.7
    Cost
    $0.012/task
    Audit
    Measured configuration
    GPT-5.6 Luna (Non-reasoning)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    28.7 Latent points at this cost
    Open cost source →
  • Sonnet 5 (max)

    Anthropic

    Prefer DeepSeek V4 Pro (max). It costs less and leads by 2.3 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    89.8
    Cost
    $2.29/task
    Audit
    Measured configuration
    Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    35.9 Latent points at this cost
    Note
    AA Intelligence Index workload at standard $3/$15 list pricing (retrieved 2026-08-17).
    Open cost source →View The Latent Score →
  • GPT-5.6 Terra (low)

    OpenAI

    Better buy: GPT-5.6 Luna (high). GPT-5.6 Luna (high) is +21.8 Latent points and 77% cheaper.

    Latent Score
    47.5
    Cost
    $0.094/task
    Audit
    Measured configuration
    GPT-5.6 Terra (low)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    42.2 Latent points at this cost
    Open cost source →
  • GPT-5.6 Sol (non-reasoning)

    OpenAI

    Better buy: GPT-5.6 Luna (high). GPT-5.6 Luna (high) is +21.0 Latent points and 91% cheaper.

    Latent Score
    48.3
    Cost
    $0.24/task
    Audit
    Measured configuration
    GPT-5.6 Sol (Non-reasoning)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    43.7 Latent points at this cost
    Open cost source →
  • GPT-5.6 Terra (non-reasoning)

    OpenAI

    Better buy: GPT-5.6 Luna (low). GPT-5.6 Luna (low) is +1.4 Latent points and 91% cheaper.

    Latent Score
    32.0
    Cost
    $0.10/task
    Audit
    Measured configuration
    GPT-5.6 Terra (Non-reasoning)
    Workload
    Artificial Analysis Intelligence Index v4.1.1
    Cost retrieved
    2026-08-17
    Gap to frontier
    57.9 Latent points at this cost
    Open cost source →

Not ranked2 of 37 models

  • GLM-5.3No Artificial Analysis cost measurement107.7score-cost/task
  • Qwen3.8-27BNo Artificial Analysis cost measurement72score-cost/task

37 published configurations across 15 models. Observed cost per task by Artificial Analysis (Intelligence Index v4.1.1 workload, retrieved 2026-08-17). Models without a published effort knob appear once.

Best value, ranked · observed cost

Every configuration, best value first

No scalar value score decides this order. Configurations clearing the Latent ≥ 80 floor rank ahead of those below it; within each group, by distance from the cost-capability frontier (0 = no better trade-off exists at its price); ties go to the cheaper configuration.

#ConfigurationLatent ScoreCost/taskGap to frontierVerdict
1GPT-5.6 Luna (max)OpenAI88.0$0.0470Best value: the cheapest configuration clearing the capability floor
2DeepSeek V4 Pro (max)DeepSeek92.1$0.250Step-up option: unbeatable at its price
3Muse Spark 1.2 (xhigh)Meta109.5$0.400Step-up option: unbeatable at its price
4Opus 5 (high)Anthropic118.9$1.230Step-up option: unbeatable at its price
5Opus 5 (xhigh)Anthropic123.2$1.800Step-up option: unbeatable at its price
6Opus 5 (max)Anthropic125.9$2.340Maximum intelligence: the strongest brain on this board. Worth the multiple only for the most complex reasoning work the benchmarks reward.
7Fable 5 (max)Anthropic122.3$3.14−3.6Prefer Opus 5 (xhigh); pay up only for the hardest reasoning (gap within uncertainty)
8Gemini 3.7 Flash (medium)Google88.5$0.26−5.1Prefer DeepSeek V4 Pro (max); pay up only for the hardest reasoning (gap within uncertainty)
9Kimi K3 (max)Moonshot AI108.7$0.84−7.0Prefer Muse Spark 1.2 (xhigh); pay up only for the hardest reasoning (gap within uncertainty)
10GPT-5.6 Sol (max)OpenAI111.5$1.23−7.4Better buy: Opus 5 (high). +7.4 points at the same cost
11Opus 5 (medium)Anthropic106.7$0.72−7.7Prefer Muse Spark 1.2 (xhigh); pay up only for the hardest reasoning (gap within uncertainty)
12Gemini 3.7 Flash (high)Google98.4$0.40−11.1Prefer Muse Spark 1.2 (xhigh); pay up only for the hardest reasoning (gap within uncertainty)
13GPT-5.6 Sol (xhigh)OpenAI103.9$0.81−11.5Prefer Muse Spark 1.2 (xhigh); pay up only for the hardest reasoning (gap within uncertainty)
14Grok 4.6 (high)xAI103.4$0.84−12.3Prefer Muse Spark 1.2 (xhigh); pay up only for the hardest reasoning (gap within uncertainty)
15GPT-5.6 Sol (high)OpenAI97.3$0.55−14.9Better buy: Muse Spark 1.2 (xhigh). +12.2 points, 27% cheaper
16GPT-5.6 Terra (max)OpenAI96.1$0.51−15.4Better buy: Muse Spark 1.2 (xhigh). +13.4 points, 22% cheaper
17GPT-5.6 Sol (medium)OpenAI91.0$0.37−15.6Prefer DeepSeek V4 Pro (max); pay up only for the hardest reasoning (gap within uncertainty)
18Qwen3.8-MaxAlibaba101.7$1.13−16.5Better buy: Muse Spark 1.2 (xhigh). +7.8 points, 65% cheaper
19GPT-5.6 Terra (xhigh)OpenAI82.3$0.31−17.8Better buy: GPT-5.6 Luna (max). +5.7 points, 85% cheaper
20Opus 5 (low)Anthropic83.4$0.43−26.7Better buy: GPT-5.6 Luna (max). +4.6 points, 89% cheaper
21Sonnet 5 (max)Anthropic89.8$2.29−35.9Prefer DeepSeek V4 Pro (max); pay up only for the hardest reasoning (gap within uncertainty)
22GPT-5.6 Luna (low)OpenAI33.4$0.0090Unbeatable at its price, but below the Latent ≥ 80 floor
23GPT-5.6 Luna (medium)OpenAI45.4$0.0110Unbeatable at its price, but below the Latent ≥ 80 floor
24GPT-5.6 Luna (high)OpenAI69.3$0.0220Unbeatable at its price, but below the Latent ≥ 80 floor
25GPT-5.6 Luna (xhigh)OpenAI79.9$0.0320Unbeatable at its price, but below the Latent ≥ 80 floor
26Gemini 3.7 Flash (low)Google79.5$0.16−11.5Prefer GPT-5.6 Luna (xhigh); pay up only for the hardest reasoning (gap within uncertainty)
27DeepSeek V4 Flash (max)DeepSeek73.0$0.11−17.1Better buy: GPT-5.6 Luna (xhigh). +6.9 points, 71% cheaper
28GPT-5.6 Sol (low)OpenAI74.0$0.23−17.9Better buy: GPT-5.6 Luna (xhigh). +5.9 points, 86% cheaper
29GPT-5.6 Terra (high)OpenAI73.1$0.22−18.7Better buy: GPT-5.6 Luna (xhigh). +6.8 points, 85% cheaper
30Kimi K3 (low)Moonshot AI67.7$0.24−24.3Prefer GPT-5.6 Luna (high); pay up only for the hardest reasoning (gap within uncertainty)
31GPT-5.6 Terra (medium)OpenAI62.8$0.12−27.5Better buy: GPT-5.6 Luna (high). +6.5 points, 82% cheaper
32GPT-5.6 Luna (non-reasoning)OpenAI19.7$0.012−28.7Better buy: GPT-5.6 Luna (low). +13.7 points, 25% cheaper
33GPT-5.6 Terra (low)OpenAI47.5$0.094−42.2Better buy: GPT-5.6 Luna (high). +21.8 points, 77% cheaper
34GPT-5.6 Sol (non-reasoning)OpenAI48.3$0.24−43.7Better buy: GPT-5.6 Luna (high). +21.0 points, 91% cheaper
35GPT-5.6 Terra (non-reasoning)OpenAI32.0$0.10−57.9Better buy: GPT-5.6 Luna (low). +1.4 points, 91% cheaper

A low rank is not a ban: the most capable configurations are unbeatable at what they do. The ranking answers "is there a better way to spend your money?"; the verdict says when there isn't. Provisional models are never ranked.

Secondary lens · listed price, fixed basket

Compare by listed token prices

Same purchasing rules, different definition of cost: 1,000,000 input + 250,000 output tokens at listed first-party rates. Reasoning-token consumption is invisible here. That is what the observed-cost view above adds.

  1. DeepSeek V4 Flash

    DeepSeek

    Latent Score
    73.0
    Cost
    $0.21 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $0.14/M in · $0.28/M out
    Price updated
    2026-08-17
    Latent 95% interval
    67.3 to 81.2
    Note
    Cache-miss list price for deepseek-v4-flash (0731 checkpoint).
    Open pricing source →View The Latent Score →
  2. +$0.29 basket · +15.0 Latent points · 2.4× the cost

    GPT-5.6 Luna

    OpenAI

    Best bang for buck

    Latent Score
    88.0
    Cost
    $0.50 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $0.2/M in · $1.2/M out
    Price updated
    2026-08-17
    Latent 95% interval
    83.2 to 95.0
    Note
    Standard API pricing.
    Open pricing source →View The Latent Score →
  3. +$1.19 basket · +10.4 Latent points · 3.4× the cost

    Gemini 3.7 Flash

    Google

    Latent Score
    98.4
    Cost
    $1.69 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $0.75/M in · $3.75/M out
    Price updated
    2026-08-16
    Latent 95% interval
    87.0 to 110.8
    Note
    Introductory rate through 2026-12-31; standard rate is $1.50/$7.50 from 2027-01-01.
    Disclosed rate scenario
    Standard rate from 2027-01-01 (introductory rate expires): basket $3.38
    Open pricing source →View The Latent Score →
  4. +$0.63 basket · +11.1 Latent points · 1.4× the cost

    Muse Spark 1.2

    Meta

    Latent Score
    109.5
    Cost
    $2.31 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $1.25/M in · $4.25/M out
    Price updated
    2026-08-16
    Latent 95% interval
    98.7 to 115.7
    Note
    Standard tier; contributor tier is excluded (training-data exchange).
    Open pricing source →View The Latent Score →
  5. +$8.94 basket · +16.4 Latent points · 4.9× the cost

    Opus 5

    Anthropic

    Maximum intelligence

    Latent Score
    125.9
    Cost
    $11.25 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $5/M in · $25/M out
    Price updated
    2026-08-15
    Latent 95% interval
    119.8 to 132.1
    Note
    Standard uncached API pricing.
    Open pricing source →View The Latent Score →

Why not these models?Each has a cheaper-or-equal alternative that scores higher.

  • Fable 5

    Anthropic

    Prefer Opus 5. It costs less and leads by 3.6 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    122.3
    Cost
    $22.50 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $10/M in · $50/M out
    Price updated
    2026-08-15
    Latent 95% interval
    116.9 to 126.3
    Gap to frontier
    3.6 Latent points at this price
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
    Note
    Standard uncached API pricing.
    Open pricing source →View The Latent Score →
  • Grok 4.6

    xAI

    Prefer Muse Spark 1.2. It costs less and leads by 6.1 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    103.4
    Cost
    $3.50 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $2/M in · $6/M out
    Price updated
    2026-08-15
    Latent 95% interval
    98.2 to 111.4
    Gap to frontier
    10.4 Latent points at this price
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
    Note
    Base tier below 200k prompt tokens.
    Open pricing source →View The Latent Score →
  • Kimi K3

    Moonshot AI

    Prefer Muse Spark 1.2. It costs less and leads by 0.8 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    108.7
    Cost
    $6.75 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $3/M in · $15/M out
    Price updated
    2026-08-15
    Latent 95% interval
    102.7 to 112.8
    Gap to frontier
    11.9 Latent points at this price
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
    Note
    Hosted API cache-miss pricing.
    Open pricing source →View The Latent Score →
  • Qwen3.8-Max

    Alibaba

    Better buy: Muse Spark 1.2. Muse Spark 1.2 is +7.8 Latent points and 34% cheaper.

    Latent Score
    101.7
    Cost
    $3.50 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $2/M in · $6/M out
    Price updated
    2026-08-16
    Latent 95% interval
    95.8 to 106.2
    Gap to frontier
    12.1 Latent points at this price
    Note
    International (Singapore) Model Studio list pricing.
    Open pricing source →View The Latent Score →
  • GPT-5.6 Sol

    OpenAI

    Better buy: Opus 5. Opus 5 is +14.4 Latent points and 10% cheaper.

    Latent Score
    111.5
    Cost
    $12.50 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $5/M in · $30/M out
    Price updated
    2026-08-15
    Latent 95% interval
    106.1 to 116.6
    Gap to frontier
    14.4 Latent points at this price
    Note
    Standard tier below the long-context threshold.
    Open pricing source →View The Latent Score →
  • DeepSeek V4 Pro

    DeepSeek

    Prefer Gemini 3.7 Flash. It costs less and leads by 6.3 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    92.1
    Cost
    $2.31 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $1.32/M in · $3.96/M out
    Price updated
    2026-08-16
    Latent 95% interval
    85.8 to 100.8
    Gap to frontier
    17.4 Latent points at this price
    Robust frontier
    Within uncertainty of the frontier: no cheaper-or-equal model is more capable at ≥90% bootstrap probability.
    Note
    Peak cache-miss rate effective 2026-08-16; off-peak is half of peak.
    Disclosed rate scenario
    Off-peak rate (published; half of peak): basket $1.16
    Open pricing source →View The Latent Score →
  • GPT-5.6 Terra

    OpenAI

    Prefer Gemini 3.7 Flash. It costs less and leads by 2.3 Latent points, though the gap sits inside uncertainty. Paying more at this tier only makes sense when you need the strongest reasoning for the most complex work the benchmarks reward.

    Latent Score
    96.1
    Cost
    $5.00 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $2/M in · $12/M out
    Price updated
    2026-08-17
    Latent 95% interval
    91.7 to 101.8
    Gap to frontier
    21.4 Latent points at this price
    Note
    Standard tier below the long-context threshold.
    Open pricing source →View The Latent Score →
  • Sonnet 5

    Anthropic

    Better buy: Gemini 3.7 Flash. Gemini 3.7 Flash is +8.6 Latent points and 62% cheaper.

    Latent Score
    89.8
    Cost
    $4.50 basket
    Audit
    Basket
    1M input + 250K output
    List price
    $2/M in · $10/M out
    Price updated
    2026-08-23
    Latent 95% interval
    80.5 to 95.3
    Gap to frontier
    26.6 Latent points at this price
    Note
    Standard API pricing.
    Open pricing source →View The Latent Score →

Not ranked2 of 15 models

  • GLM-5.3No public per-token price107.7score-basket
  • Qwen3.8-27BNo public per-token price72score-basket

le-2026-v2.0 · lv-2026-v3.0

Rule-based purchasing criteria, two cost lenses

Full methodology →

Observed cost, strictly matched

The headline view uses Artificial Analysis's measured cost to complete one Intelligence Index v4.1.1 task (input, cache, reasoning, and answer tokens included). The model board accepts only the exact configuration the sheet pins; a mismatch or missing measurement means unranked, never approximated. The configuration board applies the same rules to every published reasoning effort of the ranked models. The frontier chart additionally plots estimated effort variants for models Artificial Analysis measures only once (Fable 5, Grok 4.6, Qwen3.8-Max, Muse Spark 1.2): scores derive from CursorBench, the one public benchmark that tests every reasoning effort, calibrated on the measured effort ladders; costs derive from the measured cost ladders. Each is marked as an estimate and never ranked.

Dominance before everything

A model that another ranked model beats on both dimensions (cheaper or equal in cost, at least as capable, and strictly better on one) is never recommended ahead of its dominator. Point estimates decide placement; the bootstrap decides wording. A dominance claim is presented as decisive only when the dominator's capability advantage is positive in ≥90% of Latent bootstrap replicates. Cost is deterministic, so only capability needs the test.

A declared capability floor

The Value Pick is the cheapest non-dominated model clearing Latent ≥ 80, so an extremely cheap but weak model can never headline on price alone. Declared minimum useful-capability threshold: 1.33 SD below the 100-point cohort anchor on the 100±15 scale. A versioned policy choice, re-declared whenever the scale re-anchors; not discovered from the data.

A ladder, not a value race

Models above the pick are not "worse value" in some scalar sense. They are step-up options, ordered by cost, each stating exactly what the extra spend buys in Latent points. The scalar Efficiency and Value scores published under the previous board versions remain in the released snapshots for reproducibility, but no displayed ordering uses them.

Observed cost per task measurements by Artificial Analysis (Intelligence Index v4.1.1 workload, retrieved 2026-08-17), used as cost evidence only. Latency, throughput, reliability, and human preference are not silently folded into any headline.