Model comparison

The Latent benchmark database

Published benchmark results for 31 frontier models, kept in each benchmark's native units and grouped into 7 capability domains.

These results are scattered across lab announcements, leaderboards, and PDFs, and no two of them are reported the same way. This page gathers the benchmarks worth comparing into one table, exactly as their publishers reported them.

205benchmarks tracked
92with enough coverage to compare
113composites + excluded evidence
31frontier models tracked

20 rows on this page · listed separately

Composite indices and excluded benchmarks

Composite indices bundle up benchmarks already listed above, so they sit on their own rather than in a domain table. Rows covering fewer than 2 models are kept here until more results are published. Other rows remain separate when their reviewed evidence does not meet the scoring criteria.

BenchmarkOpus 5.5AnthropicGPT-6 AstraOpenAIGemini 4 ArgonGoogleFable 5.1AnthropicSonnet 5.5AnthropicOpus 5AnthropicMiMo-V2.6-ProXiaomiFable 5AnthropicGPT-6.1 SolOpenAIGPT-6 SolOpenAIMuse Spark 1.3MetaDeepSeek V4.1-FlashDeepSeekGPT-5.6 SolOpenAIGrok 4.7xAIGrok 4.6xAIKimi K3Moonshot AIGLM-5.3Z.aiGemini 3.8 FlashGoogleGemini 3.7 FlashGoogleClaude Haiku 5.5AnthropicGPT-6 LunaOpenAIGPT-5.6 TerraOpenAIGLM 5.3 FlashZ.aiQwen3.8-MaxAlibabaMuse Spark 1.2MetaDeepSeek V4 ProDeepSeekMistral Large 4Mistral AIGPT-5.6 LunaOpenAISonnet 5AnthropicQwen3.8-27BAlibabaDeepSeek V4 FlashDeepSeekEvidenceStatus
KingBench 3Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-----77.5%-82.5%----71.25%--77.5%91.25%----62.9%-81.25%76.25%76.25%----72.5%—Insufficient distinct-provider anchors or zero spread
LifeSciBenchCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-60.3%----------59.9%------------------—Insufficient distinct-provider anchors or zero spread
LVBench (with Memory)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-------90.1%----84.2%----87.8%85.4%----85.6%-------—Insufficient distinct-provider anchors or zero spread
MathArena — ArXivMath Aug 2026 — Anthropic no toolsCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.91.2%---86.8%--------------------------—Excluded from scoring; see methodology
MathArena — ArXivMath Aug 2026 — Anthropic with toolsCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.96.9%---95.2%--------------------------—Excluded from scoring; see methodology
MathArena ApexCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-----------65.6%-------------------—No frozen calibration; retained as collected evidence.
MedChemBench (internal)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-49.3%----------47.4%------------------—Insufficient distinct-provider anchors or zero spread
MiMo Code BenchCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.------63.2%------------------------—Excluded from scoring; see methodology
MiMo Cyber BenchCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.------80.2%------------------------—Excluded from scoring; see methodology
MiMo Visual CodingCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.------72.3%------------------------—Excluded from scoring; see methodology
MineBenchCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.---2,050-2,106-1,972--1,836-2,108-1,9311,8121,8501,9171,867----1,3331,653--1,7841,720-1,601—Unresolved Pro/base model identities in source catalog; exclude until matched.
MobileWorldCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-------85.5%----76.9%----------77.8%-------SourceCitationInsufficient distinct-provider anchors or zero spread
MRCR v2 (8 needles, 512K-1M)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-96.3%----------73.8%------------------—Insufficient distinct-provider anchors or zero spread
OfficeQA — Opus 5.5 releaseCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.78.9%---76.9%--------------------------—Excluded from scoring; see methodology
OpenLM Arena+ — Coding (Elo)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.------1,560------------------------—Excluded from scoring; see methodology
OpenLM Arena+ — Overall (Elo)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.------1,507------------------------—Excluded from scoring; see methodology
OpenScore String Quartets (1 - OMR-NED)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-0.84----------0.19------------------—Insufficient distinct-provider anchors or zero spread
OSWorld 2.0Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-----31.43%----32%-62.6%-----------17.9%------—Insufficient distinct-provider anchors or zero spread
OSWorld 2.0 — Sep 10 tasks, Anthropic partial creditCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.81.8%---80.1%--------------------------—Excluded from scoring; see methodology
OSWorld 2.0 — Sep 10 tasks, Anthropic strict passCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.48.7%---43.5%--------------------------—Excluded from scoring; see methodology

Reading this table

Native units, nothing adjusted

Units

Cells show each benchmark's native scale: accuracy percentages, Elo ratings, or points. Nothing is rescaled or converted, so compare along a row. Two different rows are two different scales.

Best in row

The strongest published result in each row is highlighted. A blank cell means that model has no published result on that benchmark, which is not the same as a zero.

What gets listed

A benchmark earns a row when its result is published against a named model version with a stated metric, and when enough models have run it for the comparison to mean something. Numbers we cannot trace back to a source are left out.

Provenance

Results are current as of 2026-10-10. Vendor-reported head-to-head rows are named as such (e.g. “Anthropic H2H”) and carry their caveats in the row notes.

Evidence downloads: manifest · dataset JSON · results CSV · citations CSV