Model comparison

The Latent benchmark database

Published benchmark results for 15 frontier models, kept in each benchmark's native units and grouped into 7 capability domains.

These results are scattered across lab announcements, leaderboards, and PDFs, and no two of them are reported the same way. This page gathers the benchmarks worth comparing into one table, exactly as their publishers reported them.

84benchmarks tracked
73with enough coverage to compare
11composites + thin coverage
15frontier models tracked

11 benchmarks in this domain

Knowledge & reasoning

Academic knowledge, mathematics, abstract reasoning, factuality, and long-context reasoning.

BenchmarkOpus 5AnthropicFable 5AnthropicGPT-5.6 SolOpenAIMuse Spark 1.2MetaKimi K3Moonshot AIGLM-5.3Z.aiGrok 4.6xAIQwen3.8-MaxAlibabaGemini 3.7 FlashGoogleGPT-5.6 TerraOpenAIDeepSeek V4 ProDeepSeekSonnet 5AnthropicGPT-5.6 LunaOpenAIDeepSeek V4 FlashDeepSeekQwen3.8-27BAlibabaEvidence
AA-LCR75.7%70%73.7%83.3%74.7%76.33%75%74.3%80%79.7%75.3%77%78.3%66%77.3%SourceCitation
AA-Omniscience31.2740.1521.727.218.4214.330.5-26.50.1-16.45-10.28-16-9.98
ARC-AGI-290.4%89.2%92.5%-60.4%-67.1%--83.9%--59.6%61.4%-SourceCitation
FrontierMath v2 Tier 473.17%87.8%82.9%-39.02%--46.3%-68.3%-29.27%61%24.4%-
GPQA Diamond93.2%92.6%94.1%90.4%93.5%91.72%94.9%92.6%94.5%92.9%92.8%91.1%92.3%91%90.5%SourceCitation
Humanity's Last Exam (no tools)52.6%53.3%47.2%45.5%44.3%42.26%42.9%43.6%47.9%42.9%41%41.3%39.5%37%33.9%
LiveBench80.1%83%81%78%79.2%-78%78.5%78.8%77.9%-76%73.6%74.2%59%SourceCitation
MathArena - ArXivMath Jun 202680.95%83.67%86.73%-72.11%--------42.86%-
MathArena - BrokenArXiv Jun 202690.74%47.84%67.28%-51.85%--------16.67%-
MMLU Pro - Vals91.6%91.5%89.1%88.28%87.97%-89.4%88.6%-86.66%-87.55%86.04%86.21%-
SimpleQA56.7%68.3%71.6%---53%--43.1%57.9%-41.7%--SourceCitation

19 benchmarks in this domain

Coding & software engineering

Agentic terminal work, issue resolution, long-horizon engineering, competitive coding, and autonomous ML engineering.

BenchmarkOpus 5AnthropicFable 5AnthropicGPT-5.6 SolOpenAIMuse Spark 1.2MetaKimi K3Moonshot AIGLM-5.3Z.aiGrok 4.6xAIQwen3.8-MaxAlibabaGemini 3.7 FlashGoogleGPT-5.6 TerraOpenAIDeepSeek V4 ProDeepSeekSonnet 5AnthropicGPT-5.6 LunaOpenAIDeepSeek V4 FlashDeepSeekQwen3.8-27BAlibabaEvidence
APEX-SWE (Pass@1, Terminus-2)54.67%58.8%41.2%-48%-56.4%----46.4%-46.9%-
Code Migration †57.5%55.1%52.9%29.95%16.1%-44.57%--36.45%41.54%44.39%44.55%38.63%-SourceCitation
CursorBench v3.270%70.5%67.2%-60.8%-69.9%-61.6%64.9%-61.5%61.1%--
DeepSWE v1.174%70%73%59.3%69%66.9%65.9%56.6%65.3%69.6%62.7%54%67.2%53%42.2%
FrontierCode v1.1 ExtendedOpus is medium effort; Grok is high; Cognition labels Kimi effort as none63.6%63.6%60.6%-58.2%-61.3%--55.8%--55.1%31.7%-
FrontierCode v1.1 Main - Anthropic H2HCognition FrontierCode 1.1 Main / BenchLM; claude-code best effort; Cognition FrontierCode 1.1 Main / BenchLM; claude-code best effort; Cognition FrontierCode 1.1 Main / BenchLM; claude-code best effort53.4%53.5%47.5%-44.2%-48%-43.6%-17.6%42.7%---
FrontierSWE-86.6%71.3%-81.2%78.1%-73.5%-------SourceCitation
KingBench 377.5%82.5%71.25%76.25%77.5%91.25%-81.25%-62.9%76.25%--72.5%-
LiveCodeBench89.03%89.8%82.6%-87.19%-88.22%87.85%-85.93%87.53%82.43%-91.6%90.3%SourceCitation
NL2Repo----58%58%-55.9%--61.1%--54.2%42.3%
PostTrainBench35.04%41.8%36.2%-32%39.8%---------
ProgramBench82.3%76.8%77.6%-62.77%19%-10.5%-0.5%-72.1%9%--SourceCitation
SciCode56%60%56%56.4%58.7%56.48%53.6%52.9%56.8%53.9%49.2%53.6%52.5%50%44.7%SourceCitation
SWE-bench Pro79.2%80%64.6%----67.7%-63.4%55.4%63.2%62.7%52.6%61.7%SourceCitation
SWE-bench Verified - Vals97%95%96.2%86.6%93.4%95.4%95.6%85.6%-75.2%96.4%79.6%93%88.8%-
SWE-Marathon v1.150%33.1%42.5%-48.1%42.5%31.9%--32.5%--24.4%--
Terminal-Bench 2.186.7%84.6%89.5%82.9%88.3%88.2%88.4%86.6%85.8%87.4%78.7%80.5%84.7%82.7%79.78%
Terminal-Bench v3.042.7%34.1%34.6%-17.4%28.3%26%-14.9%20.8%-14.6%14.32%--
Vision2Web (Avg. Frontend/Webpage/etc.)-70.5%62.1%----69%------62.9%

13 benchmarks in this domain

Agentic tool & computer use

Tool orchestration, MCP servers, business-workflow automation, browsing, and GUI computer use.

BenchmarkOpus 5AnthropicFable 5AnthropicGPT-5.6 SolOpenAIMuse Spark 1.2MetaKimi K3Moonshot AIGLM-5.3Z.aiGrok 4.6xAIQwen3.8-MaxAlibabaGemini 3.7 FlashGoogleGPT-5.6 TerraOpenAIDeepSeek V4 ProDeepSeekSonnet 5AnthropicGPT-5.6 LunaOpenAIDeepSeek V4 FlashDeepSeekQwen3.8-27BAlibabaEvidence
Agents' Last Exam-40.5%30.6%-28.3%28.5%-27%26.3%28%25.7%-30.3%25.2%20.4%SourceCitation
APEX-Agents Mean Criteria Passed60.6%59.2%56.7%-55.4%-57.5%----48.5%-51.6%-
AutomationBench - Anthropic H2HAnthropic system card §8.11.6; Zapier private held-out, max effort; Anthropic system card §8.11.6; Zapier private held-out, max effort; Anthropic system card §8.11.6; Zapier private held-out, max effort26%17.4%18.1%--------13.5%14.9%--SourceCitation
AutomationBench v1.0.6Public 600-task v1.0.6; private Opus 26.0 belongs to the separate H2H row; Public 600-task v1.0.6; private Opus 26.0 belongs to the separate H2H row; Public 600-task v1.0.6; private Opus 26.0 belongs to the separate H2H row; Public 600-task v1.0.6; private Opus 26.0 belongs to the separate H2H row50.3%46.2%45.8%-46.7%48.2%-39.8%30.44%37.17%43.2%--25.1%-
BrowseComp90.8%88%90.4%-91.2%----87.5%83.4%84.7%83.3%73.2%-SourceCitation
Humanity's Last Exam w/ tools64.7%63.9%64.5%-59.8%62.5%-56.2%--60%57.4%48.9%45.1%-
MCP Atlas †Muse uses the benchmark provider harness at xhigh effort85.8%84.7%83.6%90.3%84.2%-----73.6%----
OSWorld 2.0 - Anthropic H2H70.6%66.1%62.6%-58.3%--46.7%-50.2%--45.6%--
OSWorld-VerifiedOpus is a self-reported Ouroboros v6.87.0 run with max acting effort; Opus is a self-reported Ouroboros v6.87.0 run with max acting effort; Anthropic Sonnet 5 system card §8.10.2; pass@1 avg 5 runs, max effort; Opus is a self-reported Ouroboros v6.87.0 run with max acting effort; Anthropic Sonnet 5 system card §8.10.2; pass@1 avg 5 runs, max effort; Opus is a self-reported Ouroboros v6.87.0 run with max acting effort; Anthropic Sonnet 5 system card §8.10.2; pass@1 avg 5 runs, max effort90.69%85%83.2%-84.8%--86.1%---81.2%--84.3%SourceCitation
SkillsBench †60.4%70.9%73.5%53.04%-47.51%-70.2%-60.65%-46.48%60.45%50.67%-
Toolathlon-VerifiedAlternate Fable official-service avg@3 was 74.70% on 2026-08-14; canonical cell retains the max-labeled 77.90% snapshot; Alternate Fable official-service avg@3 was 74.70% on 2026-08-14; canonical cell retains the max-labeled 77.90% snapshot; Alternate Fable official-service avg@3 was 74.70% on 2026-08-14; canonical cell retains the max-labeled 77.90% snapshot; Alternate Fable official-service avg@3 was 74.70% on 2026-08-14; canonical cell retains the max-labeled 77.90% snapshot80.6%77.9%74.9%75.9%76.5%73%-72.5%--74.1%71.6%-70.3%-SourceCitation
WebArena-Verified-71.3%69.7%----66.8%------64.8%SourceCitation
τ²-Bench Telecom-98.5%85.1%-80.63%----86.3%96.2%----

17 benchmarks in this domain

Professional & real-world work

Economically valuable knowledge work: legal, finance, tax, medical, enterprise documents, and occupational workflows.

BenchmarkOpus 5AnthropicFable 5AnthropicGPT-5.6 SolOpenAIMuse Spark 1.2MetaKimi K3Moonshot AIGLM-5.3Z.aiGrok 4.6xAIQwen3.8-MaxAlibabaGemini 3.7 FlashGoogleGPT-5.6 TerraOpenAIDeepSeek V4 ProDeepSeekSonnet 5AnthropicGPT-5.6 LunaOpenAIDeepSeek V4 FlashDeepSeekQwen3.8-27BAlibabaEvidence
AA-Briefcase Overall Elo17201574150213581541-157714201132--1384-1286-
CorpFin v2 - Vals73.2%71.8%64.4%70.94%71.56%--65.85%-65.31%65.42%67.95%64.22%61.85%-SourceCitation
CoWorkBench-75.9%71.5%----74.8%--66.3%---70.7%SourceCitation
Excel Modeling Benchmark - Vals73.6%73.7%72.3%56.98%66.4%-62.57%60.07%71.14%67.61%52.8%66.32%67.1%56.98%-
Finance Agent v2 - Vals58.6%56.3%53.8%60.6%54.4%-53.68%50.6%59.55%52.42%50.4%53.9%55%49.5%-
GDPval-AA v2 (Elo)186117411728162816821769175317371525157815901595158115591546
Harvey LAB (Vals)6.7%11.3%2.5%25.42%10.83%-15.8%10.4%8.75%0.42%7.5%5%1.25%8.3%-SourceCitation
Harvey Legal Agent Benchmark - held-outAnthropic system card §8.11.3; Harvey held-out all-pass rate, max effort; Anthropic system card §8.11.3; Harvey held-out all-pass rate, max effort; Anthropic system card §8.11.3; Harvey held-out all-pass rate, max effort11.7%13.3%2.5%--------5.8%---
JobBench-57.4%45.4%-54.3%--53.4%------33.4%SourceCitation
Legal Research Bench - Vals55.3%49.5%48.1%43.75%44.23%-48.08%47.6%34.62%40.87%40.87%41.83%36.54%30.29%-
LegalBench - Vals87%88.6%87%85.26%86.02%-86.31%83.61%-85.11%82.36%83.92%84.03%77.71%-SourceCitation
MedCode - Vals63.6%56.1%44%49.35%48.88%-44.71%40.67%-43.41%42.47%47.54%42.39%41.41%-SourceCitation
MedScribe - Vals91%88.5%85.2%90.06%87.96%-86.53%84.95%-82.87%80.17%76.05%84.39%80.36%-SourceCitation
MortgageTax - Vals72.1%68.9%67.3%65.42%66.34%-64.19%63.99%-67.33%-70.03%67.29%--
OfficeQA Pro †Grok is the xAI model-card High-effort result; Grok is the xAI model-card High-effort result; Anthropic system card §8.11.1; OfficeQA Pro exact-match, internal harness, max effort; Grok is the xAI model-card High-effort result; Anthropic system card §8.11.1; OfficeQA Pro exact-match, internal harness, max effort; Grok is the xAI model-card High-effort result; Anthropic system card §8.11.1; OfficeQA Pro exact-match, internal harness, max effort66.9%69.9%63.2%-63.3%-63.2%----59.4%---
Public Benefits Bench - Vals76.9%70.4%66.5%68.47%68.27%-66.85%67.12%-62.38%62.92%66.03%61.16%--SourceCitation
TaxEval v2 - Vals75.1%76.9%74.8%80.38%75.72%-71.1%75.55%-76.17%73.06%75.63%76.17%70.69%-SourceCitation

7 benchmarks in this domain

Multimodal & vision

Chart, image, video, and handwriting understanding.

BenchmarkOpus 5AnthropicFable 5AnthropicGPT-5.6 SolOpenAIMuse Spark 1.2MetaKimi K3Moonshot AIGLM-5.3Z.aiGrok 4.6xAIQwen3.8-MaxAlibabaGemini 3.7 FlashGoogleGPT-5.6 TerraOpenAIDeepSeek V4 ProDeepSeekSonnet 5AnthropicGPT-5.6 LunaOpenAIDeepSeek V4 FlashDeepSeekQwen3.8-27BAlibabaEvidence
BabyVision (with CI)-90.5%88.9%-85.7%--91.3%------85.6%
CharXiv (with CI / RQ)-93.5%89.1%-91.3%--93.5%88.7%--88.3%--90.2%
ERQA-70%70%----77.8%----62.42%-65.5%SourceCitation
LVBench (with Memory)-90.1%84.2%----85.6%85.4%------
MMMU-Pro84.7%81.2%83%-81.6%--82.3%-80.7%-77.3%78.4%--SourceCitation
PerceptionBench-57.2%59.7%-58.5%--63.5%-------SourceCitation
SAGE - Vals49.4%51.9%52.6%47.66%54.26%-28.9%51.25%-47%-48.92%44.22%--SourceCitation

2 benchmarks in this domain

Cybersecurity

Vulnerability analysis and exploit development capability.

BenchmarkOpus 5AnthropicFable 5AnthropicGPT-5.6 SolOpenAIMuse Spark 1.2MetaKimi K3Moonshot AIGLM-5.3Z.aiGrok 4.6xAIQwen3.8-MaxAlibabaGemini 3.7 FlashGoogleGPT-5.6 TerraOpenAIDeepSeek V4 ProDeepSeekSonnet 5AnthropicGPT-5.6 LunaOpenAIDeepSeek V4 FlashDeepSeekQwen3.8-27BAlibabaEvidence
CyberGym-83.8%83.6%-80%84.5%79.7%78.5%-81.8%83.3%52.7%77.9%76.7%-
ExploitBench70%78%76.5%-32.2%54.4%-28.8%-52.9%--33.2%--

4 benchmarks in this domain

Preference & communication

Arena-style human preference, emotional intelligence, creative writing, and design preference.

BenchmarkOpus 5AnthropicFable 5AnthropicGPT-5.6 SolOpenAIMuse Spark 1.2MetaKimi K3Moonshot AIGLM-5.3Z.aiGrok 4.6xAIQwen3.8-MaxAlibabaGemini 3.7 FlashGoogleGPT-5.6 TerraOpenAIDeepSeek V4 ProDeepSeekSonnet 5AnthropicGPT-5.6 LunaOpenAIDeepSeek V4 FlashDeepSeekQwen3.8-27BAlibabaEvidence
Creative Writing v32429.52064.62091.8-2340.4----1927.6--18991438.4-SourceCitation
Design Arena (Elo)14071400137913731453--1388-1280-12971283--SourceCitation
EQ-Bench 4138513401250-1339----1234--1156--SourceCitation
LMArena (Chatbot Arena) - Text14891506148114991489-14641491149014641465146114501435-

11 rows · listed separately

Composite indices and thin coverage

Composite indices bundle up benchmarks already listed above, so they sit on their own rather than in a domain table. Rows covering fewer than 4 models are kept here until more results are published.

BenchmarkOpus 5AnthropicFable 5AnthropicGPT-5.6 SolOpenAIMuse Spark 1.2MetaKimi K3Moonshot AIGLM-5.3Z.aiGrok 4.6xAIQwen3.8-MaxAlibabaGemini 3.7 FlashGoogleGPT-5.6 TerraOpenAIDeepSeek V4 ProDeepSeekSonnet 5AnthropicGPT-5.6 LunaOpenAIDeepSeek V4 FlashDeepSeekQwen3.8-27BAlibabaEvidenceStatus
AA Intelligence Index v4.1.1 (score)636261576060615856575355525252Composite · aggregates other rows
AIIQ Composite IQ134134136-122----132--129--SourceCitationComposite · aggregates other rows
EQ-Bench 3-2049.7--------1570.2--1491.3-Coverage 3 < 4
Frontier-Bench v0.1 - Anthropic H2H43.3%33.7%34.4%------------SourceCitationCoverage 3 < 4
MathArena - ArXivLean Jun 202631.25%14.58%37.5%------------Coverage 3 < 4
MathArena - ArXivMath May 2026-87.5%--61.67%--------55.83%-Coverage 3 < 4
MobileWorld-85.5%76.9%----77.8%-------SourceCitationCoverage 3 < 4
PaperBench (Replication Score)-88.8%90.5%----93%-------Coverage 3 < 4
QwenReactBench (Elo)-17701564----1724-------Coverage 3 < 4
Vals Index74.8%75.1%73.1%71.88%74.7%-71.1%66.12%59.31%56.53%66.25%59.61%59.88%53.57%-SourceCitationComposite · aggregates other rows
Vals Multimodal Index73.9%74.2%72.2%69.8%73.42%--65.39%-65.07%-68.83%69.06%--SourceCitationComposite · aggregates other rows

Reading this table

Native units, nothing adjusted

Units

Cells show each benchmark's native scale: accuracy percentages, Elo ratings, or points. Nothing is rescaled or converted, so compare along a row. Two different rows are two different scales.

Best in row

The strongest published result in each row is highlighted. A blank cell means that model has no published result on that benchmark, which is not the same as a zero.

What gets listed

A benchmark earns a row when its result is published against a named model version with a stated metric, and when enough models have run it for the comparison to mean something. Numbers we cannot trace back to a source are left out.

Provenance

Results are current as of 2026-08-23. Vendor-reported head-to-head rows are named as such (e.g. “Anthropic H2H”) and carry their caveats in the row notes.

Evidence downloads: manifest · dataset JSON · results CSV · citations CSV