Model comparison
The Latent benchmark database
Published benchmark results for 31 frontier models, kept in each benchmark's native units and grouped into 7 capability domains.
These results are scattered across lab announcements, leaderboards, and PDFs, and no two of them are reported the same way. This page gathers the benchmarks worth comparing into one table, exactly as their publishers reported them.
5 benchmarks on this page
Multimodal & vision
Chart, image, video, and handwriting understanding.
| Benchmark | Opus 5.5Anthropic | GPT-6 AstraOpenAI | Gemini 4 ArgonGoogle | Fable 5.1Anthropic | Sonnet 5.5Anthropic | Opus 5Anthropic | MiMo-V2.6-ProXiaomi | Fable 5Anthropic | GPT-6.1 SolOpenAI | GPT-6 SolOpenAI | Muse Spark 1.3Meta | DeepSeek V4.1-FlashDeepSeek | GPT-5.6 SolOpenAI | Grok 4.7xAI | Grok 4.6xAI | Kimi K3Moonshot AI | GLM-5.3Z.ai | Gemini 3.8 FlashGoogle | Gemini 3.7 FlashGoogle | Claude Haiku 5.5Anthropic | GPT-6 LunaOpenAI | GPT-5.6 TerraOpenAI | GLM 5.3 FlashZ.ai | Qwen3.8-MaxAlibaba | Muse Spark 1.2Meta | DeepSeek V4 ProDeepSeek | Mistral Large 4Mistral AI | GPT-5.6 LunaOpenAI | Sonnet 5Anthropic | Qwen3.8-27BAlibaba | DeepSeek V4 FlashDeepSeek | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| PerceptionBenchCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | - | - | - | - | - | - | 57.2% | - | - | - | - | 59.7% | - | - | 58.5% | - | - | - | - | - | - | - | 63.5% | - | - | - | - | - | - | - | SourceCitation |
| Roboflow Vision Evals (six-task mean, high tier)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | 88.53% | - | 82.52% | - | - | - | - | - | - | 77.73% | - | 81% | - | - | - | - | 86.52% | 86.22% | - | - | 74.25% | - | 85% | 80.02% | - | - | 76.05% | - | - | - | — |
| SAGE - ValsCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | 46.37% | - | 48.53% | - | 49.43% | - | 51.89% | - | - | - | - | 52.56% | - | 28.9% | 54.26% | - | 35.06% | 49.23% | - | - | 47% | - | 51.25% | 47.66% | - | - | 44.22% | 48.92% | 52.4% | - | SourceCitation |
| VoxelBench (text-to-voxel, Glicko-2)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | 2,714 | - | 2,245 | - | 2,315 | - | 2,149 | - | - | - | - | 2,284 | - | 2,099 | 1,932 | 1,786 | 1,811 | 1,854 | - | - | 1,963 | - | 1,847 | - | - | - | 1,846 | 1,570 | - | 1,436 | — |
| ZeroBenchCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | - | - | - | - | 26% | - | 24% | - | - | - | - | 30% | - | 17% | 23% | - | - | - | - | - | 19% | - | 24% | - | - | - | 21% | 13% | - | - | — |
3 benchmarks on this page
Cybersecurity
Vulnerability analysis and exploit development capability.
| Benchmark | Opus 5.5Anthropic | GPT-6 AstraOpenAI | Gemini 4 ArgonGoogle | Fable 5.1Anthropic | Sonnet 5.5Anthropic | Opus 5Anthropic | MiMo-V2.6-ProXiaomi | Fable 5Anthropic | GPT-6.1 SolOpenAI | GPT-6 SolOpenAI | Muse Spark 1.3Meta | DeepSeek V4.1-FlashDeepSeek | GPT-5.6 SolOpenAI | Grok 4.7xAI | Grok 4.6xAI | Kimi K3Moonshot AI | GLM-5.3Z.ai | Gemini 3.8 FlashGoogle | Gemini 3.7 FlashGoogle | Claude Haiku 5.5Anthropic | GPT-6 LunaOpenAI | GPT-5.6 TerraOpenAI | GLM 5.3 FlashZ.ai | Qwen3.8-MaxAlibaba | Muse Spark 1.2Meta | DeepSeek V4 ProDeepSeek | Mistral Large 4Mistral AI | GPT-5.6 LunaOpenAI | Sonnet 5Anthropic | Qwen3.8-27BAlibaba | DeepSeek V4 FlashDeepSeek | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CWE-benchCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | - | - | - | - | - | - | 47.8% | - | - | - | - | 44.2% | - | 38.2% | 22.5% | 31.1% | - | 44% | - | - | - | - | 37.5% | 30.3% | - | - | - | - | - | 30.4% | — |
| CyberGymCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | - | - | - | - | - | 94% | 83.8% | - | - | - | 88.1% | 83.6% | - | 79.7% | 80% | 84.5% | - | - | - | - | 81.8% | - | 78.5% | - | 83.3% | - | 77.9% | 52.7% | - | 76.7% | — |
| SRE-Bench (pass@1)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | 88% | - | - | - | 12.5% | - | - | - | - | - | - | 55.9% | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | — |
4 benchmarks on this page
Preference & communication
Arena-style human preference, emotional intelligence, creative writing, and design preference.
| Benchmark | Opus 5.5Anthropic | GPT-6 AstraOpenAI | Gemini 4 ArgonGoogle | Fable 5.1Anthropic | Sonnet 5.5Anthropic | Opus 5Anthropic | MiMo-V2.6-ProXiaomi | Fable 5Anthropic | GPT-6.1 SolOpenAI | GPT-6 SolOpenAI | Muse Spark 1.3Meta | DeepSeek V4.1-FlashDeepSeek | GPT-5.6 SolOpenAI | Grok 4.7xAI | Grok 4.6xAI | Kimi K3Moonshot AI | GLM-5.3Z.ai | Gemini 3.8 FlashGoogle | Gemini 3.7 FlashGoogle | Claude Haiku 5.5Anthropic | GPT-6 LunaOpenAI | GPT-5.6 TerraOpenAI | GLM 5.3 FlashZ.ai | Qwen3.8-MaxAlibaba | Muse Spark 1.2Meta | DeepSeek V4 ProDeepSeek | Mistral Large 4Mistral AI | GPT-5.6 LunaOpenAI | Sonnet 5Anthropic | Qwen3.8-27BAlibaba | DeepSeek V4 FlashDeepSeek | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Creative Writing v3Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | 2,163.9 | - | 2,152.7 | - | 2,116.1 | - | 1,933.2 | - | - | - | - | 1,964.1 | - | - | 2,070.8 | 2,062.4 | - | 1,725.7 | - | - | 1,850 | - | 1,842 | 1,835.4 | - | - | 1,826.6 | 1,787.6 | 1,669.1 | 1,438.3 | SourceCitation |
| Design Arena (Elo)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | - | - | - | - | 1,407 | - | 1,400 | - | - | - | - | 1,379 | - | - | 1,453 | - | - | - | - | - | 1,280 | - | 1,388 | 1,373 | - | - | 1,283 | 1,297 | - | - | SourceCitation |
| EQ-Bench 4Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | - | - | - | - | 1,385 | - | 1,340.4 | - | - | - | - | 1,250.3 | - | - | 1,339.3 | - | - | - | - | - | 1,234 | - | - | - | - | - | 1,156.3 | 1,236 | - | - | SourceCitation |
| LMArena (Chatbot Arena) - TextCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | - | - | 1,504.21 | - | 1,493.33 | - | 1,507.16 | - | - | - | - | 1,483.15 | - | 1,461.15 | 1,488.67 | 1,482.04 | 1,494 | 1,490 | - | - | 1,466.42 | - | 1,479.56 | 1,499 | 1,459.61 | - | 1,452.6 | 1,462.37 | 1,435.89 | 1,435 | — |
8 rows on this page · listed separately
Composite indices and excluded benchmarks
Composite indices bundle up benchmarks already listed above, so they sit on their own rather than in a domain table. Rows covering fewer than 2 models are kept here until more results are published. Other rows remain separate when their reviewed evidence does not meet the scoring criteria.
| Benchmark | Opus 5.5Anthropic | GPT-6 AstraOpenAI | Gemini 4 ArgonGoogle | Fable 5.1Anthropic | Sonnet 5.5Anthropic | Opus 5Anthropic | MiMo-V2.6-ProXiaomi | Fable 5Anthropic | GPT-6.1 SolOpenAI | GPT-6 SolOpenAI | Muse Spark 1.3Meta | DeepSeek V4.1-FlashDeepSeek | GPT-5.6 SolOpenAI | Grok 4.7xAI | Grok 4.6xAI | Kimi K3Moonshot AI | GLM-5.3Z.ai | Gemini 3.8 FlashGoogle | Gemini 3.7 FlashGoogle | Claude Haiku 5.5Anthropic | GPT-6 LunaOpenAI | GPT-5.6 TerraOpenAI | GLM 5.3 FlashZ.ai | Qwen3.8-MaxAlibaba | Muse Spark 1.2Meta | DeepSeek V4 ProDeepSeek | Mistral Large 4Mistral AI | GPT-5.6 LunaOpenAI | Sonnet 5Anthropic | Qwen3.8-27BAlibaba | DeepSeek V4 FlashDeepSeek | Evidence | Status |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AA Coding Agent Index v1.4 (score)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | 67 | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | — | Composite · aggregates other rows |
| AA Intelligence Index v4.1.1 (score)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | 61.2 | - | 66 | - | 63 | - | 62 | - | - | 61 | - | 61 | - | 61 | 60 | 60 | 59 | 56 | - | - | 57 | - | 58 | 57 | 53 | - | 52 | 55 | 52 | 52 | — | Composite · aggregates other rows |
| AA Intelligence Index v4.3.2 (score)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | 57.62 | - | 52.56 | - | - | - | 46.32 | - | 51.83 | 47.53 | - | - | - | - | - | - | - | - | - | 43.4 | 38.12 | - | 41.81 | - | - | - | 38.38 | - | - | - | - | — | Composite · aggregates other rows |
| AA-Briefcase v1.1 (Elo)Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | 1,807.05 | - | 1,488.47 | - | 1,811 | - | 1,521.82 | - | 1,557.24 | 1,482.78 | - | - | - | 1,643.78 | - | - | - | - | - | 1,577.1 | 1,336.38 | - | 1,454.36 | - | - | - | 1,392.59 | - | - | - | - | — | Excluded from scoring; see methodology |
| AA-LCRCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | 74.3% | - | 80% | - | 75.7% | - | 70% | - | - | 79% | - | 73.7% | - | 75% | 74.7% | 76.33% | 82% | 80% | - | - | 79.7% | - | 74.3% | 83.3% | 75.3% | - | 78.3% | 77% | 77.3% | 66% | SourceCitation | Insufficient distinct-provider anchors or zero spread |
| AA-LCR v1.1Current values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | 84.67% | 80.67% | 79.67% | 85.33% | - | - | 86.33% | - | 84% | 83.67% | - | - | - | 77% | - | - | - | - | - | 82.67% | 83.33% | - | 80% | - | - | - | 81.33% | - | - | - | - | — | No frozen calibration; retained as collected evidence. |
| Agent Arena overall IPSCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | 0.13 | - | 0.15 | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | - | — | Excluded from scoring; see methodology |
| AIIQ Composite IQCurrent values follow the best-published-reasoning-mode policy. Individual source, configuration and admission decisions are disclosed on /methodology#model-evidence. Collected values without a matching review remain visible but do not contribute. | - | - | - | - | - | 134 | - | 134 | - | - | - | - | 136 | - | - | 122 | - | - | - | - | - | 132 | - | - | - | - | - | 129 | - | - | - | SourceCitation | Composite · aggregates other rows |
No benchmark matches your filters. Try a shorter query or clear the domain filter.
Reading this table
Native units, nothing adjusted
Units
Cells show each benchmark's native scale: accuracy percentages, Elo ratings, or points. Nothing is rescaled or converted, so compare along a row. Two different rows are two different scales.
Best in row
The strongest published result in each row is highlighted. A blank cell means that model has no published result on that benchmark, which is not the same as a zero.
What gets listed
A benchmark earns a row when its result is published against a named model version with a stated metric, and when enough models have run it for the comparison to mean something. Numbers we cannot trace back to a source are left out.
Provenance
Results are current as of 2026-10-10. Vendor-reported head-to-head rows are named as such (e.g. “Anthropic H2H”) and carry their caveats in the row notes.
Evidence downloads: manifest · dataset JSON · results CSV · citations CSV