Latent Value & Efficiency
Compare model capability and cost.
A shortlist based on current scores and documented costs. The Value Pick is the cheapest option on the point-score frontier above our declared capability floor. Observed cost comes from Artificial Analysis, including reasoning tokens. Small score gaps do not establish a reliable capability advantage.
Scores combine the best published reasoning settings across benchmarks; cost describes one measured configuration and workload. This does not promise the same capability at that cost. Test shortlisted models on your workload before choosing.
Value Pick · starting point
GPT-5.6 Luna
The cheapest frontier model above the declared score floor. New evidence or a different workload can change the comparison.
- Latent Score
- 81.3
- Observed cost / task
- $0.047
- vs Sonnet 5
- 98% cheaper
+0.0 Latent points
Capability × observed cost
The Latent Frontier
Up and to the left is better. Filled red points are the buying ladder: no cheaper option has a higher current point score.
- Buying ladder
- Dominated
Opus 5.5, Gemini 4 Argon, Sonnet 5.5, MiMo-V2.6-Pro, GPT-6.1 Sol, GPT-6 Sol, DeepSeek V4.1-Flash, Grok 4.7, GLM-5.3, Claude Haiku 5.5, GPT-6 Luna, GLM 5.3 Flash, Mistral Large 4, Qwen3.8-27B has no Artificial Analysis cost measurement and cannot be plotted. Cost data: Artificial Analysis Intelligence Index v4.1.1, retrieved 2026-08-17.
Observed cost · Artificial Analysis
Compare higher-scoring options
The ladder follows current point scores and documented costs, cheapest first. The Value Pick is the cheapest frontier model clearing the declared Latent ≥ 80 floor. Small score differences do not establish reliable superiority.
GPT-5.6 Luna
OpenAI
Value Pick
- Latent Score
- 81.3
- Cost
- $0.047/task
Audit
- Measured configuration
- GPT-5.6 Luna (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 73.0 to 90.1
- Gap to frontier
- 0.0 Latent points at this cost
- Note
- AA now measures Luna at max effort: the configuration the sheet pins (retrieved 2026-08-17).
+$0.20/task · +4.9 Latent points · 5.3× the cost
DeepSeek V4 Pro
DeepSeek
- Latent Score
- 86.2
- Cost
- $0.25/task
Audit
- Measured configuration
- DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 77.1 to 95.3
- Gap to frontier
- 0.0 Latent points at this cost
- Note
- AA measures the 0813 API snapshot at Max effort: the configuration the sheet pins.
+$0.15/task · +15.2 Latent points · 1.6× the cost
Gemini 3.7 Flash
Google
- Latent Score
- 101.4
- Cost
- $0.40/task
Audit
- Measured configuration
- Gemini 3.7 Flash (high)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 91.2 to 111.5
- Gap to frontier
- 0.0 Latent points at this cost
+$0.15/task · +12.8 Latent points · 1.4× the cost
Muse Spark 1.3
Meta
- Latent Score
- 114.2
- Cost
- $0.55/task
Audit
- Measured configuration
- Muse Spark 1.3 (xhigh)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 102.3 to 125.8
- Gap to frontier
- 0.0 Latent points at this cost
- Note
- AA Intelligence Index v4.1.1 at xhigh; retrieved 2026-09-03. Max is limited preview without public pricing and is not the sheet pin.
+$1.12/task · +25.0 Latent points · 3× the cost
GPT-6 Astra
OpenAI
Highest current score
- Latent Score
- 139.2
- Cost
- $1.67/task
Audit
- Measured configuration
- GPT-6 Astra (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 128.0 to 150.2
- Gap to frontier
- 0.0 Latent points at this cost
- Note
- AA Intelligence Index v4.1.1 at max effort; retrieved 2026-09-03 (launch day).
Why not these models?Each has an alternative with no higher cost and no lower current score.
GPT-5.6 Sol
OpenAI
Compare Muse Spark 1.3. It costs less and leads by 2.3 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 111.9
- Cost
- $1.23/task
Audit
- Measured configuration
- GPT-5.6 Sol (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 103.8 to 119.6
- Gap to frontier
- 2.3 Latent points at this cost
Fable 5.1
Anthropic
Compare GPT-6 Astra. It costs less and leads by 4.5 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 134.7
- Cost
- $3.69/task
Audit
- Measured configuration
- Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Opus 5 Fallback)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 125.6 to 143.8
- Gap to frontier
- 4.5 Latent points at this cost
- Note
- AA Intelligence Index v4.1.1 at max effort with documented Opus 5 fallback.
GPT-5.6 Terra
OpenAI
Compare Gemini 3.7 Flash. It costs less and leads by 6.3 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 95.1
- Cost
- $0.51/task
Audit
- Measured configuration
- GPT-5.6 Terra (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 88.8 to 101.3
- Gap to frontier
- 6.3 Latent points at this cost
- Note
- AA now measures Terra at max effort: the configuration the sheet pins (retrieved 2026-08-17).
Grok 4.6
xAI
Compare Muse Spark 1.3. It costs less and leads by 8.5 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 105.7
- Cost
- $0.84/task
Audit
- Measured configuration
- Grok 4.6 (high)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 97.1 to 115.5
- Gap to frontier
- 8.5 Latent points at this cost
Muse Spark 1.2
Meta
Compare Gemini 3.7 Flash. It leads by 9.2 Latent points at the same cost. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 92.2
- Cost
- $0.40/task
Audit
- Measured configuration
- Muse Spark 1.2 (xhigh)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 83.7 to 100.7
- Gap to frontier
- 9.2 Latent points at this cost
Opus 5
Anthropic
Compare GPT-6 Astra. It costs less and leads by 10.0 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 129.2
- Cost
- $2.34/task
Audit
- Measured configuration
- Claude Opus 5 (Adaptive Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 123.6 to 134.8
- Gap to frontier
- 10.0 Latent points at this cost
Kimi K3
Moonshot AI
Compare Muse Spark 1.3. It costs less and leads by 10.2 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 104.0
- Cost
- $0.84/task
Audit
- Measured configuration
- Kimi K3 (max)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 95.2 to 112.5
- Gap to frontier
- 10.2 Latent points at this cost
Gemini 3.8 Flash
Google
Compare Muse Spark 1.3. It costs less and leads by 11.4 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 102.8
- Cost
- $0.58/task
Audit
- Measured configuration
- Gemini 3.8 Flash (high)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 90.4 to 113.4
- Gap to frontier
- 11.4 Latent points at this cost
- Note
- AA Intelligence Index v4.1.1 at high effort; retrieved 2026-09-03. Token list price matches 3.7 Flash; observed cost is higher because the model uses more reasoning tokens.
DeepSeek V4 Flash
DeepSeek
Compare GPT-5.6 Luna. It costs less and leads by 14.2 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 67.1
- Cost
- $0.11/task
Audit
- Measured configuration
- DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 59.5 to 73.3
- Gap to frontier
- 14.2 Latent points at this cost
- Note
- AA measures Flash 0731 at Max effort: the configuration the sheet pins.
Fable 5
Anthropic
Compare GPT-6 Astra. It costs less and leads by 19.2 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 120.0
- Cost
- $3.14/task
Audit
- Measured configuration
- Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 112.9 to 126.8
- Gap to frontier
- 19.2 Latent points at this cost
- Note
- AA's harness used the documented Opus 4.8 fallback path; disclosed by AA.
Qwen3.8-Max
Alibaba
Compare Gemini 3.7 Flash. It costs less and leads by 8.3 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 93.1
- Cost
- $1.13/task
Audit
- Measured configuration
- Qwen3.8 Max
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 83.8 to 103.2
- Gap to frontier
- 21.1 Latent points at this cost
Sonnet 5
Anthropic
Compare GPT-5.6 Luna. It costs less and leads by 0.0 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 81.3
- Cost
- $2.29/task
Audit
- Measured configuration
- Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
- Workload
- Artificial Analysis Intelligence Index v4.1.1
- Cost retrieved
- 2026-08-17
- Score stability range
- 74.3 to 88.5
- Gap to frontier
- 57.9 Latent points at this cost
- Note
- AA Intelligence Index workload at standard $3/$15 list pricing (retrieved 2026-08-17).
Not ranked14 of 31 models
- Opus 5.5Insufficient comparable evidence or no configuration-matched cost.144.2score-cost/task
- Gemini 4 ArgonInsufficient comparable evidence or no configuration-matched cost.134.8score-cost/task
- Sonnet 5.5Insufficient comparable evidence or no configuration-matched cost.132.5score-cost/task
- MiMo-V2.6-ProInsufficient comparable evidence or no configuration-matched cost.125.8score-cost/task
- GPT-6.1 SolInsufficient comparable evidence or no configuration-matched cost.118.2score-cost/task
- GPT-6 SolInsufficient comparable evidence or no configuration-matched cost.117.5score-cost/task
- DeepSeek V4.1-FlashInsufficient comparable evidence or no configuration-matched cost.113score-cost/task
- Grok 4.7Insufficient comparable evidence or no configuration-matched cost.111.3score-cost/task
- GLM-5.3Insufficient comparable evidence or no configuration-matched cost.103.8score-cost/task
- Claude Haiku 5.5Insufficient comparable evidence or no configuration-matched cost.100.5score-cost/task
- GPT-6 LunaInsufficient comparable evidence or no configuration-matched cost.97.3score-cost/task
- GLM 5.3 FlashInsufficient comparable evidence or no configuration-matched cost.94.4score-cost/task
- Mistral Large 4Insufficient comparable evidence or no configuration-matched cost.83score-cost/task
- Qwen3.8-27BInsufficient comparable evidence or no configuration-matched cost.70.5score-cost/task
Secondary lens · listed price, fixed basket
Compare by listed token prices
Same purchasing rules, different definition of cost: 1,000,000 input + 250,000 output tokens at listed first-party rates. Reasoning-token consumption is invisible here. That is what the observed-cost view above adds.
DeepSeek V4 Flash
DeepSeek
- Latent Score
- 67.1
- Cost
- $0.21 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $0.14/M in · $0.28/M out
- Price updated
- 2026-08-17
- Score stability range
- 59.5 to 73.3
- Gap to frontier
- 0.0 Latent points at this price
- Note
- Cache-miss list price for deepseek-v4-flash (0731 checkpoint).
+$0.01 basket · +33.4 Latent points · 1.1× the cost
Claude Haiku 5.5
Anthropic
Value Pick
- Latent Score
- 100.5
- Cost
- $0.23 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $0.1/M in · $0.5/M out
- Price updated
- 2026-10-07
- Score stability range
- 83.8 to 117.2
- Gap to frontier
- 0.0 Latent points at this price
- Note
- Listed rates as published by the evaluator on 2026-10-07. Cache reads are 0.01 per million.
+$0.43 basket · +25.3 Latent points · 2.9× the cost
MiMo-V2.6-Pro
Xiaomi
- Latent Score
- 125.8
- Cost
- $0.65 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $0.435/M in · $0.87/M out
- Price updated
- 2026-09-22
- Score stability range
- 112.0 to 138.3
- Gap to frontier
- 0.0 Latent points at this price
- Note
- Standard API cache-miss USD pricing per million tokens; cache hits $0.0036/M. UltraSpeed is separately priced.
+$3.85 basket · +9.0 Latent points · 6.9× the cost
Gemini 4 Argon
Google
- Latent Score
- 134.8
- Cost
- $4.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $10/M out
- Price updated
- 2026-09-30
- Score stability range
- 123.4 to 146.9
- Gap to frontier
- 0.0 Latent points at this price
- Note
- Listed rates as published by the evaluator on 2026-09-30. Cache reads are 0.1 per million.
+$4.50 basket · +9.4 Latent points · 2× the cost
Opus 5.5
Anthropic
Highest current score
- Latent Score
- 144.2
- Cost
- $9.00 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $4/M in · $20/M out
- Price updated
- 2026-09-22
- Score stability range
- 133.3 to 155.2
- Gap to frontier
- 0.0 Latent points at this price
- Note
- Standard API rates; cached reads $0.20/M. Benchmark evidence includes disclosed production fallback routing.
Why not these models?Each has an alternative with no higher cost and no lower current score.
Sonnet 5.5
Anthropic
Compare Gemini 4 Argon. It leads by 2.3 Latent points at the same cost. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 132.5
- Cost
- $4.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $10/M out
- Price updated
- 2026-09-28
- Score stability range
- 118.6 to 147.0
- Gap to frontier
- 2.3 Latent points at this price
- Note
- Standard API rates; cache reads $0.20/M and writes $2.50/M. Benchmark evidence discloses production cyber fallback to Sonnet 5.
GPT-6 Luna
OpenAI
Compare Claude Haiku 5.5. It leads by 3.2 Latent points at the same cost. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 97.3
- Cost
- $0.23 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $0.1/M in · $0.5/M out
- Price updated
- 2026-09-22
- Score stability range
- 80.8 to 116.9
- Gap to frontier
- 3.2 Latent points at this price
- Note
- Standard API rates up to 272K input tokens; above that, input/cache rates double and output is 1.5x. Cache reads are 10% of input, writes 1.25x. Fast and regional premiums are separate.
GPT-6 Astra
OpenAI
Compare Opus 5.5. It costs less and leads by 5.0 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 139.2
- Cost
- $22.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $10/M in · $50/M out
- Price updated
- 2026-09-03
- Score stability range
- 128.0 to 150.2
- Gap to frontier
- 5.0 Latent points at this price
- Note
- Standard API pricing at launch; cached input $1/M, cache write $12.50/M.
GLM 5.3 Flash
Z.ai
Compare Claude Haiku 5.5. It costs less and leads by 6.1 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 94.4
- Cost
- $0.28 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $0.15/M in · $0.5/M out
- Price updated
- 2026-09-30
- Score stability range
- 80.3 to 108.1
- Gap to frontier
- 6.1 Latent points at this price
- Note
- Listed rates as published by the evaluator on 2026-09-30. Cache reads are 0.026 per million.
Fable 5.1
Anthropic
Compare Gemini 4 Argon. It costs less and leads by 0.1 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 134.7
- Cost
- $22.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $10/M in · $50/M out
- Price updated
- 2026-09-01
- Score stability range
- 125.6 to 143.8
- Gap to frontier
- 9.5 Latent points at this price
- Note
- Standard uncached API pricing; cache reads $0.25/M (75% below Fable 5).
Muse Spark 1.3
Meta
Compare MiMo-V2.6-Pro. It costs less and leads by 11.6 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 114.2
- Cost
- $2.31 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $1.25/M in · $4.25/M out
- Price updated
- 2026-09-03
- Score stability range
- 102.3 to 125.8
- Gap to frontier
- 11.6 Latent points at this price
- Note
- Unchanged from 1.2: $1.25/$4.25 per 1M; cache hits $0.15. Sheet pins xhigh (generally available). Max is limited preview without public token pricing.
Grok 4.7
xAI
Compare MiMo-V2.6-Pro. It costs less and leads by 14.5 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 111.3
- Cost
- $3.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $6/M out
- Price updated
- 2026-09-21
- Score stability range
- 94.2 to 129.6
- Gap to frontier
- 14.5 Latent points at this price
- Note
- Standard API starting price; Fast runs at twice the price. Long-context rates may differ.
Opus 5
Anthropic
Compare Gemini 4 Argon. It costs less and leads by 5.6 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 129.2
- Cost
- $11.25 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $5/M in · $25/M out
- Price updated
- 2026-08-15
- Score stability range
- 123.6 to 134.8
- Gap to frontier
- 15.0 Latent points at this price
- Note
- Standard uncached API pricing.
GPT-6.1 Sol
OpenAI
Compare MiMo-V2.6-Pro. It costs less and leads by 7.6 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 118.2
- Cost
- $4.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $10/M out
- Price updated
- 2026-09-30
- Score stability range
- 100.3 to 136.0
- Gap to frontier
- 16.6 Latent points at this price
- Note
- Standard rates match GPT-6 Sol. Cache reads are 0.10 per million, 5% of input, a deeper discount than the GPT-6 Sol tier.
GPT-6 Sol
OpenAI
Compare MiMo-V2.6-Pro. It costs less and leads by 8.3 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 117.5
- Cost
- $4.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $10/M out
- Price updated
- 2026-09-22
- Score stability range
- 103.7 to 132.8
- Gap to frontier
- 17.3 Latent points at this price
- Note
- Standard API rates up to 272K input tokens; above that, input/cache rates double and output is 1.5x. Cache reads are 10% of input, writes 1.25x. Fast and regional premiums are separate.
GPT-5.6 Luna
OpenAI
Compare Claude Haiku 5.5. It costs less and leads by 19.2 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 81.3
- Cost
- $0.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $0.2/M in · $1.2/M out
- Price updated
- 2026-08-17
- Score stability range
- 73.0 to 90.1
- Gap to frontier
- 19.2 Latent points at this price
- Note
- Standard API pricing.
Grok 4.6
xAI
Compare MiMo-V2.6-Pro. It costs less and leads by 20.1 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 105.7
- Cost
- $3.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $6/M out
- Price updated
- 2026-08-15
- Score stability range
- 97.1 to 115.5
- Gap to frontier
- 20.1 Latent points at this price
- Note
- Base tier below 200k prompt tokens.
Gemini 3.8 Flash
Google
Compare MiMo-V2.6-Pro. It costs less and leads by 23.0 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 102.8
- Cost
- $1.69 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $0.75/M in · $3.75/M out
- Price updated
- 2026-09-03
- Score stability range
- 90.4 to 113.4
- Gap to frontier
- 23.0 Latent points at this price
- Note
- Introductory rate through 2026-12-31; standard rate is $1.50/$7.50 from 2027-01-01. Same list price as Gemini 3.7 Flash.
- Disclosed rate scenario
- Standard rate from 2027-01-01 (introductory rate expires): basket $3.38
Fable 5
Anthropic
Compare MiMo-V2.6-Pro. It costs less and leads by 5.8 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 120.0
- Cost
- $22.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $10/M in · $50/M out
- Price updated
- 2026-08-15
- Score stability range
- 112.9 to 126.8
- Gap to frontier
- 24.2 Latent points at this price
- Note
- Standard uncached API pricing.
Gemini 3.7 Flash
Google
Compare MiMo-V2.6-Pro. It costs less and leads by 24.4 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 101.4
- Cost
- $1.69 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $0.75/M in · $3.75/M out
- Price updated
- 2026-08-16
- Score stability range
- 91.2 to 111.5
- Gap to frontier
- 24.4 Latent points at this price
- Note
- Introductory rate through 2026-12-31; standard rate is $1.50/$7.50 from 2027-01-01.
- Disclosed rate scenario
- Standard rate from 2027-01-01 (introductory rate expires): basket $3.38
Kimi K3
Moonshot AI
Compare MiMo-V2.6-Pro. It costs less and leads by 21.8 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 104.0
- Cost
- $6.75 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $3/M in · $15/M out
- Price updated
- 2026-08-15
- Score stability range
- 95.2 to 112.5
- Gap to frontier
- 30.8 Latent points at this price
- Note
- Hosted API cache-miss pricing.
GPT-5.6 Sol
OpenAI
Compare MiMo-V2.6-Pro. It costs less and leads by 13.9 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 111.9
- Cost
- $12.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $5/M in · $30/M out
- Price updated
- 2026-08-15
- Score stability range
- 103.8 to 119.6
- Gap to frontier
- 32.3 Latent points at this price
- Note
- Standard tier below the long-context threshold.
Qwen3.8-Max
Alibaba
Compare Claude Haiku 5.5. It costs less and leads by 7.4 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 93.1
- Cost
- $3.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $6/M out
- Price updated
- 2026-08-16
- Score stability range
- 83.8 to 103.2
- Gap to frontier
- 32.7 Latent points at this price
- Note
- International (Singapore) Model Studio list pricing.
Muse Spark 1.2
Meta
Compare Claude Haiku 5.5. It costs less and leads by 8.3 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 92.2
- Cost
- $2.31 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $1.25/M in · $4.25/M out
- Price updated
- 2026-08-16
- Score stability range
- 83.7 to 100.7
- Gap to frontier
- 33.6 Latent points at this price
- Note
- Standard tier; contributor tier is excluded (training-data exchange).
DeepSeek V4 Pro
DeepSeek
Compare Claude Haiku 5.5. It costs less and leads by 14.3 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 86.2
- Cost
- $2.31 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $1.32/M in · $3.96/M out
- Price updated
- 2026-08-16
- Score stability range
- 77.1 to 95.3
- Gap to frontier
- 39.6 Latent points at this price
- Note
- Peak cache-miss rate effective 2026-08-16; off-peak is half of peak.
- Disclosed rate scenario
- Off-peak rate (published; half of peak): basket $1.16
GPT-5.6 Terra
OpenAI
Compare Claude Haiku 5.5. It costs less and leads by 5.4 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 95.1
- Cost
- $5.00 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $12/M out
- Price updated
- 2026-08-17
- Score stability range
- 88.8 to 101.3
- Gap to frontier
- 39.7 Latent points at this price
- Note
- Standard tier below the long-context threshold.
Mistral Large 4
Mistral AI
Compare Claude Haiku 5.5. It costs less and leads by 17.5 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 83.0
- Cost
- $2.40 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $1.36/M in · $4.18/M out
- Price updated
- 2026-10-06
- Score stability range
- 65.7 to 100.7
- Gap to frontier
- 42.8 Latent points at this price
- Note
- Listed rates as published by the evaluator on 2026-10-06. Cache reads are 0.14 per million.
Sonnet 5
Anthropic
Compare Claude Haiku 5.5. It costs less and leads by 19.2 Latent points. This point-score comparison does not establish a reliable capability advantage; evaluate both on the tasks and settings you plan to use.
- Latent Score
- 81.3
- Cost
- $4.50 basket
Audit
- Basket
- 1M input + 250K output
- List price
- $2/M in · $10/M out
- Price updated
- 2026-08-23
- Score stability range
- 74.3 to 88.5
- Gap to frontier
- 53.5 Latent points at this price
- Note
- Standard API pricing.
Not ranked3 of 31 models
- DeepSeek V4.1-FlashInsufficient comparable evidence or no configuration-matched cost.113score-basket
- GLM-5.3Insufficient comparable evidence or no configuration-matched cost.103.8score-basket
- Qwen3.8-27BInsufficient comparable evidence or no configuration-matched cost.70.5score-basket
le-2026-v3.0 · lv-2026-v4.0
Rule-based purchasing criteria, two cost lenses
What the cost measures
The headline view uses Artificial Analysis's measured cost to complete one Intelligence Index v4.1.1 task (input, cache, reasoning, and answer tokens included). Cost uses the recorded configuration; a missing or mismatched measurement means no cost ranking. Capability combines results across reasoning settings, so the score and measured cost do not describe one end-to-end operating point. Lower-effort score estimates from the earlier methodology are not included.
How the shortlist is ordered
The frontier contains models with no cheaper-or-equal alternative that has an equal-or-higher current score, with at least one strict advantage. This is a numerical comparison, not a finding that one model will perform better on your tasks. Placement follows point scores. Probability-of-superiority claims are not assessed in this methodology; consult each model’s interval and reporting sensitivity on the score page.
A declared capability floor
The Value Pick is the cheapest non-dominated model clearing Latent ≥ 80, so an extremely cheap but weak model can never headline on price alone. Declared capability floor of 80 on this version’s frozen 100±15 reference scale; a purchasing policy, not an empirical optimum.
A ladder, not a value race
Models above the pick are not "worse value" in some scalar sense. They are step-up options, ordered by cost, each stating the difference in current score alongside the extra cost. These differences are estimates, not guaranteed gains. The scalar Efficiency and Value scores published under the previous board versions remain in the released snapshots for reproducibility, but no displayed ordering uses them.
Observed cost per task measurements by Artificial Analysis (Intelligence Index v4.1.1 workload, retrieved 2026-08-17), used as cost evidence only. Latency, throughput, reliability, and human preference are not silently folded into any headline.