Model comparison

The Latent benchmark database

Published benchmark results for 31 frontier models, kept in each benchmark's native units and grouped into 7 capability domains.

These results are scattered across lab announcements, leaderboards, and PDFs, and no two of them are reported the same way. This page gathers the benchmarks worth comparing into one table, exactly as their publishers reported them.

205benchmarks tracked
92with enough coverage to compare
113composites + excluded evidence
31frontier models tracked

17 benchmarks on this page

Coding & software engineering

Agentic terminal work, issue resolution, long-horizon engineering, competitive coding, and autonomous ML engineering.

BenchmarkOpus 5.5AnthropicGPT-6 AstraOpenAIGemini 4 ArgonGoogleFable 5.1AnthropicSonnet 5.5AnthropicOpus 5AnthropicMiMo-V2.6-ProXiaomiFable 5AnthropicGPT-6.1 SolOpenAIGPT-6 SolOpenAIMuse Spark 1.3MetaDeepSeek V4.1-FlashDeepSeekGPT-5.6 SolOpenAIGrok 4.7xAIGrok 4.6xAIKimi K3Moonshot AIGLM-5.3Z.aiGemini 3.8 FlashGoogleGemini 3.7 FlashGoogleClaude Haiku 5.5AnthropicGPT-6 LunaOpenAIGPT-5.6 TerraOpenAIGLM 5.3 FlashZ.aiQwen3.8-MaxAlibabaMuse Spark 1.2MetaDeepSeek V4 ProDeepSeekMistral Large 4Mistral AIGPT-5.6 LunaOpenAISonnet 5AnthropicQwen3.8-27BAlibabaDeepSeek V4 FlashDeepSeekEvidence
CursorBench v3.2Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.---73.4%-70%-70.5%----67.2%-70.8%60.8%-69.2%61.6%--64.9%-----61.1%61.5%--—
DeepSWE v1.1Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.74.2%74.12%77.9%67.4%71%73.65%71.9%69.91%-68.8%75.4%74.2%72.67%71%67.48%68.51%68.96%73.83%65.49%-66.6%69.62%-57.46%54.87%62.83%-67.19%53.85%42.2%53.32%—
Frontier-Bench v0.1 - Anthropic H2HCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-----43.3%-33.7%----34.4%------------------SourceCitation
FrontierCode v1.1 ExtendedCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.65.3%64.48%-63.6%64.4%63.63%-64.94%-60.69%--60.55%-61.31%58.19%-53.45%--56.1%55.84%-----55.06%56.18%-31.7%—
FrontierSWECurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.---57%---86.6%----71.3%--81.2%78.1%------73.5%-------SourceCitation
Internal Database Migration Tasks - OpenAICurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-63.9%-57.8%---50.3%----42.7%------------------—
LiveCodeBenchCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.---90.52%-89.03%-89.78%----82.6%-88.22%87.19%80.53%89.48%88.65%--85.93%-87.85%-87.53%--82.43%84%87.26%SourceCitation
NL2RepoCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-----------65.4%---58%58%------55.9%-61.5%---42.3%54.2%—
PostTrainBenchCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-----35.04%-41.8%----34.6%--36.6%39.8%--------------—
ProgramBench v1 (Raw Pass Rate) - ValsCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-85.42%-82.69%-82.27%------77.64%--62.77%66.25%71.9%68.66%--72.35%-39.97%-70.1%-68.29%72.07%11.17%-Source
SciCodeCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.66.9%56.48%61.81%63.08%-56.37%60.88%60%55.79%57.64%59.72%-57.75%57.75%56.48%59.49%59.03%56.6%59.84%54.98%54.63%54.98%51.62%53.24%57.41%51.04%54.17%53.59%54.28%46.64%50.35%SourceCitation
SWE-bench Verified - ValsCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-----97%-95%----96.2%-95.6%93.4%95.4%80%80.8%--95.4%-85.6%86.6%96.4%-93%79.6%86%88.8%—
Terminal-Bench 2.1Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-89.89%-91.4%-89.14%89.9%84.6%--85.77%90.6%89.51%-88.39%85.02%83.9%87.64%85.77%--88.01%-81.27%80.15%78.65%-80.9%80.52%79.78%78.65%—
Terminal-Bench 4.0Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.66.36%58.2%57.4%57.9%70.6%51.8%34.9%44.5%---31.2%37.3%38%20.3%-41.8%19.1%11.2%39.2%-21.5%-----17.3%12.4%--—
Terminal-Bench Science 0.1Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.58.7%64.6%-52.6%59.9%30%-21.4%----22.4%------------------—
Terminal-Bench v3.0Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-----42.7%-34.1%---30%34.6%-26.5%17.4%32.4%-14.9%--20.8%-----14.3%14.6%--—
Vibe Code Bench v1.1 - ValsCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-89.59%91.9%90.26%-88.4%-90.35%----80.5%-76.24%84.96%78.13%78.65%70.39%--74.59%-64.7%79.1%82.3%-77.06%81.33%64.85%74.74%—

3 benchmarks on this page

Agentic tool & computer use

Tool orchestration, MCP servers, business-workflow automation, browsing, and GUI computer use.

BenchmarkOpus 5.5AnthropicGPT-6 AstraOpenAIGemini 4 ArgonGoogleFable 5.1AnthropicSonnet 5.5AnthropicOpus 5AnthropicMiMo-V2.6-ProXiaomiFable 5AnthropicGPT-6.1 SolOpenAIGPT-6 SolOpenAIMuse Spark 1.3MetaDeepSeek V4.1-FlashDeepSeekGPT-5.6 SolOpenAIGrok 4.7xAIGrok 4.6xAIKimi K3Moonshot AIGLM-5.3Z.aiGemini 3.8 FlashGoogleGemini 3.7 FlashGoogleClaude Haiku 5.5AnthropicGPT-6 LunaOpenAIGPT-5.6 TerraOpenAIGLM 5.3 FlashZ.aiQwen3.8-MaxAlibabaMuse Spark 1.2MetaDeepSeek V4 ProDeepSeekMistral Large 4Mistral AIGPT-5.6 LunaOpenAISonnet 5AnthropicQwen3.8-27BAlibabaDeepSeek V4 FlashDeepSeekEvidence
Agents' Last ExamCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-----31.6%31.6%25.7%---31.8%30.6%--28.3%28.5%-26.3%--28%26.3%27%-12.4%-30.3%-20.4%25.2%SourceCitation
Agents' Last Exam - OpenAI H2HCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-59.3%---55.5%-48.7%-56.4%--53.6%------------------—
APEX-Agents Mean Criteria PassedCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute.-75.1%-77.1%-60.6%-59.2%----56.7%-57.5%55.4%------------48.5%-51.6%—

Reading this table

Native units, nothing adjusted

Units

Cells show each benchmark's native scale: accuracy percentages, Elo ratings, or points. Nothing is rescaled or converted, so compare along a row. Two different rows are two different scales.

Best in row

The strongest published result in each row is highlighted. A blank cell means that model has no published result on that benchmark, which is not the same as a zero.

What gets listed

A benchmark earns a row when its result is published against a named model version with a stated metric, and when enough models have run it for the comparison to mean something. Numbers we cannot trace back to a source are left out.

Provenance

Results are current as of 2026-10-10. Vendor-reported head-to-head rows are named as such (e.g. “Anthropic H2H”) and carry their caveats in the row notes.

Evidence downloads: manifest · dataset JSON · results CSV · citations CSV