Model comparison
The Latent benchmark database
Published benchmark results for 15 frontier models, kept in each benchmark's native units and grouped into 7 capability domains.
These results are scattered across lab announcements, leaderboards, and PDFs, and no two of them are reported the same way. This page gathers the benchmarks worth comparing into one table, exactly as their publishers reported them.
11 benchmarks in this domain
Knowledge & reasoning
Academic knowledge, mathematics, abstract reasoning, factuality, and long-context reasoning.
| Benchmark | Opus 5Anthropic | Fable 5Anthropic | GPT-5.6 SolOpenAI | Muse Spark 1.2Meta | Kimi K3Moonshot AI | GLM-5.3Z.ai | Grok 4.6xAI | Qwen3.8-MaxAlibaba | Gemini 3.7 FlashGoogle | GPT-5.6 TerraOpenAI | DeepSeek V4 ProDeepSeek | Sonnet 5Anthropic | GPT-5.6 LunaOpenAI | DeepSeek V4 FlashDeepSeek | Qwen3.8-27BAlibaba | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AA-LCR | 75.7% | 70% | 73.7% | 83.3% | 74.7% | 76.33% | 75% | 74.3% | 80% | 79.7% | 75.3% | 77% | 78.3% | 66% | 77.3% | SourceCitation |
| AA-Omniscience | 31.27 | 40.15 | 21.7 | 27.2 | 18.42 | 14.3 | 30.5 | - | 26.5 | 0.1 | - | 16.45 | -10.28 | -16 | -9.98 | — |
| ARC-AGI-2 | 90.4% | 89.2% | 92.5% | - | 60.4% | - | 67.1% | - | - | 83.9% | - | - | 59.6% | 61.4% | - | SourceCitation |
| FrontierMath v2 Tier 4 | 73.17% | 87.8% | 82.9% | - | 39.02% | - | - | 46.3% | - | 68.3% | - | 29.27% | 61% | 24.4% | - | — |
| GPQA Diamond | 93.2% | 92.6% | 94.1% | 90.4% | 93.5% | 91.72% | 94.9% | 92.6% | 94.5% | 92.9% | 92.8% | 91.1% | 92.3% | 91% | 90.5% | SourceCitation |
| Humanity's Last Exam (no tools) | 52.6% | 53.3% | 47.2% | 45.5% | 44.3% | 42.26% | 42.9% | 43.6% | 47.9% | 42.9% | 41% | 41.3% | 39.5% | 37% | 33.9% | — |
| LiveBench | 80.1% | 83% | 81% | 78% | 79.2% | - | 78% | 78.5% | 78.8% | 77.9% | - | 76% | 73.6% | 74.2% | 59% | SourceCitation |
| MathArena - ArXivMath Jun 2026 | 80.95% | 83.67% | 86.73% | - | 72.11% | - | - | - | - | - | - | - | - | 42.86% | - | — |
| MathArena - BrokenArXiv Jun 2026 | 90.74% | 47.84% | 67.28% | - | 51.85% | - | - | - | - | - | - | - | - | 16.67% | - | — |
| MMLU Pro - Vals | 91.6% | 91.5% | 89.1% | 88.28% | 87.97% | - | 89.4% | 88.6% | - | 86.66% | - | 87.55% | 86.04% | 86.21% | - | — |
| SimpleQA | 56.7% | 68.3% | 71.6% | - | - | - | 53% | - | - | 43.1% | 57.9% | - | 41.7% | - | - | SourceCitation |
19 benchmarks in this domain
Coding & software engineering
Agentic terminal work, issue resolution, long-horizon engineering, competitive coding, and autonomous ML engineering.
| Benchmark | Opus 5Anthropic | Fable 5Anthropic | GPT-5.6 SolOpenAI | Muse Spark 1.2Meta | Kimi K3Moonshot AI | GLM-5.3Z.ai | Grok 4.6xAI | Qwen3.8-MaxAlibaba | Gemini 3.7 FlashGoogle | GPT-5.6 TerraOpenAI | DeepSeek V4 ProDeepSeek | Sonnet 5Anthropic | GPT-5.6 LunaOpenAI | DeepSeek V4 FlashDeepSeek | Qwen3.8-27BAlibaba | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| APEX-SWE (Pass@1, Terminus-2) | 54.67% | 58.8% | 41.2% | - | 48% | - | 56.4% | - | - | - | - | 46.4% | - | 46.9% | - | — |
| Code Migration † | 57.5% | 55.1% | 52.9% | 29.95% | 16.1% | - | 44.57% | - | - | 36.45% | 41.54% | 44.39% | 44.55% | 38.63% | - | SourceCitation |
| CursorBench v3.2 | 70% | 70.5% | 67.2% | - | 60.8% | - | 69.9% | - | 61.6% | 64.9% | - | 61.5% | 61.1% | - | - | — |
| DeepSWE v1.1 | 74% | 70% | 73% | 59.3% | 69% | 66.9% | 65.9% | 56.6% | 65.3% | 69.6% | 62.7% | 54% | 67.2% | 53% | 42.2% | — |
| FrontierCode v1.1 ExtendedOpus is medium effort; Grok is high; Cognition labels Kimi effort as none | 63.6% | 63.6% | 60.6% | - | 58.2% | - | 61.3% | - | - | 55.8% | - | - | 55.1% | 31.7% | - | — |
| FrontierCode v1.1 Main - Anthropic H2HCognition FrontierCode 1.1 Main / BenchLM; claude-code best effort; Cognition FrontierCode 1.1 Main / BenchLM; claude-code best effort; Cognition FrontierCode 1.1 Main / BenchLM; claude-code best effort | 53.4% | 53.5% | 47.5% | - | 44.2% | - | 48% | - | 43.6% | - | 17.6% | 42.7% | - | - | - | — |
| FrontierSWE | - | 86.6% | 71.3% | - | 81.2% | 78.1% | - | 73.5% | - | - | - | - | - | - | - | SourceCitation |
| KingBench 3 | 77.5% | 82.5% | 71.25% | 76.25% | 77.5% | 91.25% | - | 81.25% | - | 62.9% | 76.25% | - | - | 72.5% | - | — |
| LiveCodeBench | 89.03% | 89.8% | 82.6% | - | 87.19% | - | 88.22% | 87.85% | - | 85.93% | 87.53% | 82.43% | - | 91.6% | 90.3% | SourceCitation |
| NL2Repo | - | - | - | - | 58% | 58% | - | 55.9% | - | - | 61.1% | - | - | 54.2% | 42.3% | — |
| PostTrainBench | 35.04% | 41.8% | 36.2% | - | 32% | 39.8% | - | - | - | - | - | - | - | - | - | — |
| ProgramBench | 82.3% | 76.8% | 77.6% | - | 62.77% | 19% | - | 10.5% | - | 0.5% | - | 72.1% | 9% | - | - | SourceCitation |
| SciCode | 56% | 60% | 56% | 56.4% | 58.7% | 56.48% | 53.6% | 52.9% | 56.8% | 53.9% | 49.2% | 53.6% | 52.5% | 50% | 44.7% | SourceCitation |
| SWE-bench Pro | 79.2% | 80% | 64.6% | - | - | - | - | 67.7% | - | 63.4% | 55.4% | 63.2% | 62.7% | 52.6% | 61.7% | SourceCitation |
| SWE-bench Verified - Vals | 97% | 95% | 96.2% | 86.6% | 93.4% | 95.4% | 95.6% | 85.6% | - | 75.2% | 96.4% | 79.6% | 93% | 88.8% | - | — |
| SWE-Marathon v1.1 | 50% | 33.1% | 42.5% | - | 48.1% | 42.5% | 31.9% | - | - | 32.5% | - | - | 24.4% | - | - | — |
| Terminal-Bench 2.1 | 86.7% | 84.6% | 89.5% | 82.9% | 88.3% | 88.2% | 88.4% | 86.6% | 85.8% | 87.4% | 78.7% | 80.5% | 84.7% | 82.7% | 79.78% | — |
| Terminal-Bench v3.0 | 42.7% | 34.1% | 34.6% | - | 17.4% | 28.3% | 26% | - | 14.9% | 20.8% | - | 14.6% | 14.32% | - | - | — |
| Vision2Web (Avg. Frontend/Webpage/etc.) | - | 70.5% | 62.1% | - | - | - | - | 69% | - | - | - | - | - | - | 62.9% | — |
13 benchmarks in this domain
Agentic tool & computer use
Tool orchestration, MCP servers, business-workflow automation, browsing, and GUI computer use.
| Benchmark | Opus 5Anthropic | Fable 5Anthropic | GPT-5.6 SolOpenAI | Muse Spark 1.2Meta | Kimi K3Moonshot AI | GLM-5.3Z.ai | Grok 4.6xAI | Qwen3.8-MaxAlibaba | Gemini 3.7 FlashGoogle | GPT-5.6 TerraOpenAI | DeepSeek V4 ProDeepSeek | Sonnet 5Anthropic | GPT-5.6 LunaOpenAI | DeepSeek V4 FlashDeepSeek | Qwen3.8-27BAlibaba | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Agents' Last Exam | - | 40.5% | 30.6% | - | 28.3% | 28.5% | - | 27% | 26.3% | 28% | 25.7% | - | 30.3% | 25.2% | 20.4% | SourceCitation |
| APEX-Agents Mean Criteria Passed | 60.6% | 59.2% | 56.7% | - | 55.4% | - | 57.5% | - | - | - | - | 48.5% | - | 51.6% | - | — |
| AutomationBench - Anthropic H2HAnthropic system card §8.11.6; Zapier private held-out, max effort; Anthropic system card §8.11.6; Zapier private held-out, max effort; Anthropic system card §8.11.6; Zapier private held-out, max effort | 26% | 17.4% | 18.1% | - | - | - | - | - | - | - | - | 13.5% | 14.9% | - | - | SourceCitation |
| AutomationBench v1.0.6Public 600-task v1.0.6; private Opus 26.0 belongs to the separate H2H row; Public 600-task v1.0.6; private Opus 26.0 belongs to the separate H2H row; Public 600-task v1.0.6; private Opus 26.0 belongs to the separate H2H row; Public 600-task v1.0.6; private Opus 26.0 belongs to the separate H2H row | 50.3% | 46.2% | 45.8% | - | 46.7% | 48.2% | - | 39.8% | 30.44% | 37.17% | 43.2% | - | - | 25.1% | - | — |
| BrowseComp | 90.8% | 88% | 90.4% | - | 91.2% | - | - | - | - | 87.5% | 83.4% | 84.7% | 83.3% | 73.2% | - | SourceCitation |
| Humanity's Last Exam w/ tools | 64.7% | 63.9% | 64.5% | - | 59.8% | 62.5% | - | 56.2% | - | - | 60% | 57.4% | 48.9% | 45.1% | - | — |
| MCP Atlas †Muse uses the benchmark provider harness at xhigh effort | 85.8% | 84.7% | 83.6% | 90.3% | 84.2% | - | - | - | - | - | 73.6% | - | - | - | - | — |
| OSWorld 2.0 - Anthropic H2H | 70.6% | 66.1% | 62.6% | - | 58.3% | - | - | 46.7% | - | 50.2% | - | - | 45.6% | - | - | — |
| OSWorld-VerifiedOpus is a self-reported Ouroboros v6.87.0 run with max acting effort; Opus is a self-reported Ouroboros v6.87.0 run with max acting effort; Anthropic Sonnet 5 system card §8.10.2; pass@1 avg 5 runs, max effort; Opus is a self-reported Ouroboros v6.87.0 run with max acting effort; Anthropic Sonnet 5 system card §8.10.2; pass@1 avg 5 runs, max effort; Opus is a self-reported Ouroboros v6.87.0 run with max acting effort; Anthropic Sonnet 5 system card §8.10.2; pass@1 avg 5 runs, max effort | 90.69% | 85% | 83.2% | - | 84.8% | - | - | 86.1% | - | - | - | 81.2% | - | - | 84.3% | SourceCitation |
| SkillsBench † | 60.4% | 70.9% | 73.5% | 53.04% | - | 47.51% | - | 70.2% | - | 60.65% | - | 46.48% | 60.45% | 50.67% | - | — |
| Toolathlon-VerifiedAlternate Fable official-service avg@3 was 74.70% on 2026-08-14; canonical cell retains the max-labeled 77.90% snapshot; Alternate Fable official-service avg@3 was 74.70% on 2026-08-14; canonical cell retains the max-labeled 77.90% snapshot; Alternate Fable official-service avg@3 was 74.70% on 2026-08-14; canonical cell retains the max-labeled 77.90% snapshot; Alternate Fable official-service avg@3 was 74.70% on 2026-08-14; canonical cell retains the max-labeled 77.90% snapshot | 80.6% | 77.9% | 74.9% | 75.9% | 76.5% | 73% | - | 72.5% | - | - | 74.1% | 71.6% | - | 70.3% | - | SourceCitation |
| WebArena-Verified | - | 71.3% | 69.7% | - | - | - | - | 66.8% | - | - | - | - | - | - | 64.8% | SourceCitation |
| τ²-Bench Telecom | - | 98.5% | 85.1% | - | 80.63% | - | - | - | - | 86.3% | 96.2% | - | - | - | - | — |
17 benchmarks in this domain
Professional & real-world work
Economically valuable knowledge work: legal, finance, tax, medical, enterprise documents, and occupational workflows.
| Benchmark | Opus 5Anthropic | Fable 5Anthropic | GPT-5.6 SolOpenAI | Muse Spark 1.2Meta | Kimi K3Moonshot AI | GLM-5.3Z.ai | Grok 4.6xAI | Qwen3.8-MaxAlibaba | Gemini 3.7 FlashGoogle | GPT-5.6 TerraOpenAI | DeepSeek V4 ProDeepSeek | Sonnet 5Anthropic | GPT-5.6 LunaOpenAI | DeepSeek V4 FlashDeepSeek | Qwen3.8-27BAlibaba | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AA-Briefcase Overall Elo | 1720 | 1574 | 1502 | 1358 | 1541 | - | 1577 | 1420 | 1132 | - | - | 1384 | - | 1286 | - | — |
| CorpFin v2 - Vals | 73.2% | 71.8% | 64.4% | 70.94% | 71.56% | - | - | 65.85% | - | 65.31% | 65.42% | 67.95% | 64.22% | 61.85% | - | SourceCitation |
| CoWorkBench | - | 75.9% | 71.5% | - | - | - | - | 74.8% | - | - | 66.3% | - | - | - | 70.7% | SourceCitation |
| Excel Modeling Benchmark - Vals | 73.6% | 73.7% | 72.3% | 56.98% | 66.4% | - | 62.57% | 60.07% | 71.14% | 67.61% | 52.8% | 66.32% | 67.1% | 56.98% | - | — |
| Finance Agent v2 - Vals | 58.6% | 56.3% | 53.8% | 60.6% | 54.4% | - | 53.68% | 50.6% | 59.55% | 52.42% | 50.4% | 53.9% | 55% | 49.5% | - | — |
| GDPval-AA v2 (Elo) | 1861 | 1741 | 1728 | 1628 | 1682 | 1769 | 1753 | 1737 | 1525 | 1578 | 1590 | 1595 | 1581 | 1559 | 1546 | — |
| Harvey LAB (Vals) | 6.7% | 11.3% | 2.5% | 25.42% | 10.83% | - | 15.8% | 10.4% | 8.75% | 0.42% | 7.5% | 5% | 1.25% | 8.3% | - | SourceCitation |
| Harvey Legal Agent Benchmark - held-outAnthropic system card §8.11.3; Harvey held-out all-pass rate, max effort; Anthropic system card §8.11.3; Harvey held-out all-pass rate, max effort; Anthropic system card §8.11.3; Harvey held-out all-pass rate, max effort | 11.7% | 13.3% | 2.5% | - | - | - | - | - | - | - | - | 5.8% | - | - | - | — |
| JobBench | - | 57.4% | 45.4% | - | 54.3% | - | - | 53.4% | - | - | - | - | - | - | 33.4% | SourceCitation |
| Legal Research Bench - Vals | 55.3% | 49.5% | 48.1% | 43.75% | 44.23% | - | 48.08% | 47.6% | 34.62% | 40.87% | 40.87% | 41.83% | 36.54% | 30.29% | - | — |
| LegalBench - Vals | 87% | 88.6% | 87% | 85.26% | 86.02% | - | 86.31% | 83.61% | - | 85.11% | 82.36% | 83.92% | 84.03% | 77.71% | - | SourceCitation |
| MedCode - Vals | 63.6% | 56.1% | 44% | 49.35% | 48.88% | - | 44.71% | 40.67% | - | 43.41% | 42.47% | 47.54% | 42.39% | 41.41% | - | SourceCitation |
| MedScribe - Vals | 91% | 88.5% | 85.2% | 90.06% | 87.96% | - | 86.53% | 84.95% | - | 82.87% | 80.17% | 76.05% | 84.39% | 80.36% | - | SourceCitation |
| MortgageTax - Vals | 72.1% | 68.9% | 67.3% | 65.42% | 66.34% | - | 64.19% | 63.99% | - | 67.33% | - | 70.03% | 67.29% | - | - | — |
| OfficeQA Pro †Grok is the xAI model-card High-effort result; Grok is the xAI model-card High-effort result; Anthropic system card §8.11.1; OfficeQA Pro exact-match, internal harness, max effort; Grok is the xAI model-card High-effort result; Anthropic system card §8.11.1; OfficeQA Pro exact-match, internal harness, max effort; Grok is the xAI model-card High-effort result; Anthropic system card §8.11.1; OfficeQA Pro exact-match, internal harness, max effort | 66.9% | 69.9% | 63.2% | - | 63.3% | - | 63.2% | - | - | - | - | 59.4% | - | - | - | — |
| Public Benefits Bench - Vals | 76.9% | 70.4% | 66.5% | 68.47% | 68.27% | - | 66.85% | 67.12% | - | 62.38% | 62.92% | 66.03% | 61.16% | - | - | SourceCitation |
| TaxEval v2 - Vals | 75.1% | 76.9% | 74.8% | 80.38% | 75.72% | - | 71.1% | 75.55% | - | 76.17% | 73.06% | 75.63% | 76.17% | 70.69% | - | SourceCitation |
7 benchmarks in this domain
Multimodal & vision
Chart, image, video, and handwriting understanding.
| Benchmark | Opus 5Anthropic | Fable 5Anthropic | GPT-5.6 SolOpenAI | Muse Spark 1.2Meta | Kimi K3Moonshot AI | GLM-5.3Z.ai | Grok 4.6xAI | Qwen3.8-MaxAlibaba | Gemini 3.7 FlashGoogle | GPT-5.6 TerraOpenAI | DeepSeek V4 ProDeepSeek | Sonnet 5Anthropic | GPT-5.6 LunaOpenAI | DeepSeek V4 FlashDeepSeek | Qwen3.8-27BAlibaba | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BabyVision (with CI) | - | 90.5% | 88.9% | - | 85.7% | - | - | 91.3% | - | - | - | - | - | - | 85.6% | — |
| CharXiv (with CI / RQ) | - | 93.5% | 89.1% | - | 91.3% | - | - | 93.5% | 88.7% | - | - | 88.3% | - | - | 90.2% | — |
| ERQA | - | 70% | 70% | - | - | - | - | 77.8% | - | - | - | - | 62.42% | - | 65.5% | SourceCitation |
| LVBench (with Memory) | - | 90.1% | 84.2% | - | - | - | - | 85.6% | 85.4% | - | - | - | - | - | - | — |
| MMMU-Pro | 84.7% | 81.2% | 83% | - | 81.6% | - | - | 82.3% | - | 80.7% | - | 77.3% | 78.4% | - | - | SourceCitation |
| PerceptionBench | - | 57.2% | 59.7% | - | 58.5% | - | - | 63.5% | - | - | - | - | - | - | - | SourceCitation |
| SAGE - Vals | 49.4% | 51.9% | 52.6% | 47.66% | 54.26% | - | 28.9% | 51.25% | - | 47% | - | 48.92% | 44.22% | - | - | SourceCitation |
2 benchmarks in this domain
Cybersecurity
Vulnerability analysis and exploit development capability.
| Benchmark | Opus 5Anthropic | Fable 5Anthropic | GPT-5.6 SolOpenAI | Muse Spark 1.2Meta | Kimi K3Moonshot AI | GLM-5.3Z.ai | Grok 4.6xAI | Qwen3.8-MaxAlibaba | Gemini 3.7 FlashGoogle | GPT-5.6 TerraOpenAI | DeepSeek V4 ProDeepSeek | Sonnet 5Anthropic | GPT-5.6 LunaOpenAI | DeepSeek V4 FlashDeepSeek | Qwen3.8-27BAlibaba | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CyberGym | - | 83.8% | 83.6% | - | 80% | 84.5% | 79.7% | 78.5% | - | 81.8% | 83.3% | 52.7% | 77.9% | 76.7% | - | — |
| ExploitBench | 70% | 78% | 76.5% | - | 32.2% | 54.4% | - | 28.8% | - | 52.9% | - | - | 33.2% | - | - | — |
4 benchmarks in this domain
Preference & communication
Arena-style human preference, emotional intelligence, creative writing, and design preference.
| Benchmark | Opus 5Anthropic | Fable 5Anthropic | GPT-5.6 SolOpenAI | Muse Spark 1.2Meta | Kimi K3Moonshot AI | GLM-5.3Z.ai | Grok 4.6xAI | Qwen3.8-MaxAlibaba | Gemini 3.7 FlashGoogle | GPT-5.6 TerraOpenAI | DeepSeek V4 ProDeepSeek | Sonnet 5Anthropic | GPT-5.6 LunaOpenAI | DeepSeek V4 FlashDeepSeek | Qwen3.8-27BAlibaba | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Creative Writing v3 | 2429.5 | 2064.6 | 2091.8 | - | 2340.4 | - | - | - | - | 1927.6 | - | - | 1899 | 1438.4 | - | SourceCitation |
| Design Arena (Elo) | 1407 | 1400 | 1379 | 1373 | 1453 | - | - | 1388 | - | 1280 | - | 1297 | 1283 | - | - | SourceCitation |
| EQ-Bench 4 | 1385 | 1340 | 1250 | - | 1339 | - | - | - | - | 1234 | - | - | 1156 | - | - | SourceCitation |
| LMArena (Chatbot Arena) - Text | 1489 | 1506 | 1481 | 1499 | 1489 | - | 1464 | 1491 | 1490 | 1464 | 1465 | 1461 | 1450 | 1435 | - | — |
11 rows · listed separately
Composite indices and thin coverage
Composite indices bundle up benchmarks already listed above, so they sit on their own rather than in a domain table. Rows covering fewer than 4 models are kept here until more results are published.
| Benchmark | Opus 5Anthropic | Fable 5Anthropic | GPT-5.6 SolOpenAI | Muse Spark 1.2Meta | Kimi K3Moonshot AI | GLM-5.3Z.ai | Grok 4.6xAI | Qwen3.8-MaxAlibaba | Gemini 3.7 FlashGoogle | GPT-5.6 TerraOpenAI | DeepSeek V4 ProDeepSeek | Sonnet 5Anthropic | GPT-5.6 LunaOpenAI | DeepSeek V4 FlashDeepSeek | Qwen3.8-27BAlibaba | Evidence | Status |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AA Intelligence Index v4.1.1 (score) | 63 | 62 | 61 | 57 | 60 | 60 | 61 | 58 | 56 | 57 | 53 | 55 | 52 | 52 | 52 | — | Composite · aggregates other rows |
| AIIQ Composite IQ | 134 | 134 | 136 | - | 122 | - | - | - | - | 132 | - | - | 129 | - | - | SourceCitation | Composite · aggregates other rows |
| EQ-Bench 3 | - | 2049.7 | - | - | - | - | - | - | - | - | 1570.2 | - | - | 1491.3 | - | — | Coverage 3 < 4 |
| Frontier-Bench v0.1 - Anthropic H2H | 43.3% | 33.7% | 34.4% | - | - | - | - | - | - | - | - | - | - | - | - | SourceCitation | Coverage 3 < 4 |
| MathArena - ArXivLean Jun 2026 | 31.25% | 14.58% | 37.5% | - | - | - | - | - | - | - | - | - | - | - | - | — | Coverage 3 < 4 |
| MathArena - ArXivMath May 2026 | - | 87.5% | - | - | 61.67% | - | - | - | - | - | - | - | - | 55.83% | - | — | Coverage 3 < 4 |
| MobileWorld | - | 85.5% | 76.9% | - | - | - | - | 77.8% | - | - | - | - | - | - | - | SourceCitation | Coverage 3 < 4 |
| PaperBench (Replication Score) | - | 88.8% | 90.5% | - | - | - | - | 93% | - | - | - | - | - | - | - | — | Coverage 3 < 4 |
| QwenReactBench (Elo) | - | 1770 | 1564 | - | - | - | - | 1724 | - | - | - | - | - | - | - | — | Coverage 3 < 4 |
| Vals Index | 74.8% | 75.1% | 73.1% | 71.88% | 74.7% | - | 71.1% | 66.12% | 59.31% | 56.53% | 66.25% | 59.61% | 59.88% | 53.57% | - | SourceCitation | Composite · aggregates other rows |
| Vals Multimodal Index | 73.9% | 74.2% | 72.2% | 69.8% | 73.42% | - | - | 65.39% | - | 65.07% | - | 68.83% | 69.06% | - | - | SourceCitation | Composite · aggregates other rows |
Reading this table
Native units, nothing adjusted
Units
Cells show each benchmark's native scale: accuracy percentages, Elo ratings, or points. Nothing is rescaled or converted, so compare along a row. Two different rows are two different scales.
Best in row
The strongest published result in each row is highlighted. A blank cell means that model has no published result on that benchmark, which is not the same as a zero.
What gets listed
A benchmark earns a row when its result is published against a named model version with a stated metric, and when enough models have run it for the comparison to mean something. Numbers we cannot trace back to a source are left out.
Provenance
Results are current as of 2026-08-23. Vendor-reported head-to-head rows are named as such (e.g. “Anthropic H2H”) and carry their caveats in the row notes.
Evidence downloads: manifest · dataset JSON · results CSV · citations CSV