Methodology · v1.0 · 2026-09-10
The methodology in detail.
One score summarizes demonstrated capability across published reasoning settings and seven equally weighted domains. Related benchmark variants share a family budget, and influence from one evaluator is capped.
Evidence release tli-2026-v1.0-2026-09-10-deepseek-3. Browse dated releases.
92 admitted benchmarks across 7 domains. The scale is anchored once to mean 100 and SD 15 across 18 reference models. The reference average stays fixed; 100 is neither human IQ nor the average of current models.
Evidence you can inspect
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
Seven capability domains
Each of seven capability domains has an equal, fixed importance budget. These weights define the breadth valued by the index; they are not objectively optimal or learned importance weights. Missing domain abilities are conditional estimates, explicitly marked as unmeasured; individual missing benchmark results are never invented. Alternative weights and domain groupings are sensitivity checks, not replacements chosen to favor a model.
Weight and grouping sensitivity
This release checks 35 alternatives: halving or doubling each domain’s weight, and combining every pair of domains into one equally budgeted group. The largest rank movement relative to the equal-weight diagnostic is 2 places. These checks hold the published domain estimates fixed; they do not reassign benchmarks or refit the model. Rounded domain values can differ from published tie ordering. These are sensitivity scenarios, not confidence ranges or a search for preferred weights. Published scores and ranks are unchanged.
One score and hard rank
Models are ranked by the displayed point score once they have four contributing benchmarks across two domains. Ties use stable model-name ordering. Ranks are a current ordering, not a confidence test. Scores are published immediately without a provisional label; early scores can move substantially as evidence arrives. Contributing benchmark counts and measured domain coverage appear beside each score.
Score stability range
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
Sensitivity to selective reporting
The separate reporting sensitivity range tests what happens when only the strongest or weakest standardized benchmark results are retained. It has no probability interpretation and is not combined with the score stability range. Stress-test coverage describes recovery of fuller-data scores under these deliberate masks, not real-world coverage or true intelligence. This test does not quantify the separate advantage from trying more reasoning settings.
Scores you can cite
A model’s score and interval depend only on its own accepted evidence and the frozen methodology. Other models’ additions and updates cannot change them. A methodology revision receives a new version and can change every score.
Validation and its limits
Validation includes provider and benchmark-family holdouts, retrospective sparse-to-full score recovery and selective-reporting stress tests. These test prediction and score stability within the benchmark evidence; they do not establish capability on independent real-world tasks. Prospective tracking records dated evidence snapshots for later comparison. Existing models at tracking start are baselines, not observed launches. Independent-task evaluation requires a registered protocol and outcomes excluded from scoring; no completed independent-task validation is claimed.
Value comparisons
The buying ladder compares current point scores with documented costs. Small score differences do not establish a reliable capability advantage, and interval overlap is not used as a significance test. The Value Pick is a shortlist starting point under the declared score floor, not a universal best buy. Capability combines the best published reasoning settings across benchmarks; observed cost describes one measured configuration and workload. Their combination is not a promise of that capability at that cost. Probability-of-superiority claims are not assessed.
Collected benchmarks and scoring evidence
The Contributing column on the score page shows the number of results used to calculate each score. The disclosures below show collected totals, independent family counts and reporting sensitivity. Unresolved source records, excluded composites and incompatible evaluations remain in the collected total. Related benchmarks can belong to the same family, so counts are not independent sample sizes.
The disclosures below list every collected result, whether it contributes, and its review status. An unresolved result awaits review; it has not necessarily been found incorrect or unhelpful. Collected totals alone do not establish scoring coverage or certainty.
Coverage and configuration
Use the highest published benchmark score across reasoning modes for the same model checkpoint and evaluation setup. A lower effort can win; each benchmark still contributes only once. A sole published setting with unspecified effort stays labelled as such. Retry budgets, tools, fallback models, task versions and harness differences are separate comparison questions and are not treated as reasoning modes.
This measures demonstrated capability across settings, not performance at one fixed effort or cost. We select each setting’s published aggregate, not its luckiest individual run. Models tested at more settings have more opportunities to record a high result. The main interval does not correct for that selection advantage; the chosen configuration and source are disclosed below.
A model receives a hard rank once it has 4 contributing benchmarks across 2 domains. Below that floor its score and uncertainty remain visible. No score is labelled provisional.
Cost and speed do not enter the capability score. Benchmark counts describe available evidence, not an equal number of independent experiments. More results usually improve certainty; conflicting results can increase uncertainty. A score assembled from the best mode on each benchmark is not a promise of that performance at the cost of one fixed mode.
Model evidence and uncertainty
GPT-6 Astra · 54 used / 83 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
44 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.2 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1580.2 · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); Artificial Analysis published evaluation harness
Exact checkpoint, metric and published aggregate matched to the evaluator source. Date is retrieval date; first publication date is not inferred.
- Terminal-Bench 2.1: 89.88764% · Contributes. Source · Evaluator-reported · GPT-6 Astra (high); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 74.115044% · Contributes. Source · Evaluator-reported · gpt-6-astra; xhigh; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- AA Intelligence Index v4.1.1 (score): 61.2 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 96.262626% · Contributes. Source · Evaluator-reported · GPT-6 Astra (xhigh); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 74.3% · Not used. Source · Evaluator-reported · GPT-6 Astra (max)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 53.54% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 71.7% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 56.481481% · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 54.680259% · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 39.42% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 5.42% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1562.01 · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); Artificial Analysis published evaluation harness
Exact checkpoint, metric and published aggregate matched to the evaluator source. Date is retrieval date; first publication date is not inferred.
- MedScribe ; Vals: 87.91% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 48.49% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Omniscience: 43.733333 · Contributes. Source · Evaluator-reported · GPT-6 Astra (high); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- SAGE ; Vals: 46.37% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 67.74% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Humanity's Last Exam w/ tools: 57.2% · Contributes. Source · Developer-reported · GPT-6 Astra; OpenAI comparison table; maximum across reasoning modes
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 85.417% · Contributes. Source · Archived source review · openai/gpt-6-astra · reasoning_effort=max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- FrontierMath v2 Tier 4: 97.560976% · Contributes. Source · Evaluator-reported · gpt-6-astra_medium; task version 2.0.0
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- APEX-Agents Mean Criteria Passed: 75.1% · Contributes. Source · Evaluator-reported · gpt-6-astra-max; max; loop_truncated_tools_agent
Same 240-task, 31-world APEX benchmark and mean-score metric. Current published repeated-run aggregate (959 samples); pass@1 is not substituted.
- BrowseComp: 91.5% · Contributes. Source · Developer-reported · GPT-6 Astra; OpenAI comparison table; maximum across reasoning modes
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MMMU-Pro: 86.878613% · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ARC-AGI-2: 95% · Contributes. Source · Evaluator-reported · GPT-6 Astra; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- ExploitBench: 100% · Not used.
Inherited comparator metric/scaffold equivalence to Astra Cap Percent is not established; the source flags possible contamination. Excluded pending protocol audit.
- FrontierCode v1.1 Extended: 64.48% · Contributes. Source · Evaluator-reported · GPT-6 Astra; max; codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierCode v1.1 Main ; Anthropic H2H: 53.3% · Not used.
Not admitted after the September 6 model review: the collected 53.3 for GPT-6 Astra on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AutomationBench ; Anthropic H2H: 41.4% · Not used.
Not admitted after the September 6 model review: the collected 41.4 for GPT-6 Astra on AutomationBench ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 43.092784% · Contributes. Source · Evaluator-reported · GPT-6 Astra (xhigh); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 31.714286% · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 4.0: 58.2% · Contributes. Source · Evaluator-reported · GPT-6 Astra; max; Codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 99% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 89.59% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ARC-AGI-1: 98.5% · Contributes. Source · Evaluator-reported · GPT-6 Astra; XHigh; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- ARC-AGI-3 (standard harness): 62.71% · Contributes. Source · Evaluator-reported · GPT-6 Astra; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- BenchCAD (Vision2Code subset): 95.9% · Not used. Source · Developer-reported · OpenAI full Vision2Code set; not Anthropic subset
System card identifies the OpenAI comparison as the full 17,900-file set, while this row is a modified 1000-file subset. Requires a separate protocol row.
- Terminal-Bench Science 0.1: 64.6% · Contributes. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- SRE-Bench (pass@1): 88% · Contributes. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- OSWorld 2.0 ; OpenAI H2H (offline subset): 72.6% · Contributes. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Agents' Last Exam ; OpenAI H2H: 59.3% · Contributes. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- ScreenSpot-Pro: 92.7% · Not used. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- HealthBench: 59.7% · Not used. Source · Developer-reported · GPT-6 Astra; OpenAI published HealthBench research evaluation; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- HealthBench Hard: 37.8% · Not used. Source · Developer-reported · GPT-6 Astra; OpenAI published HealthBench research evaluation; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- HealthBench Professional (length-adjusted): 63.4% · Contributes. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- GeneBench Pro: 37.1% · Not used. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- LifeSciBench: 60.3% · Not used. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- SEC-Bench Pro: 85.4% · Not used. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- ExploitGym: 42.4% · Not used. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- AA Coding Agent Index v1.4 (score): 67 · Not used.
Held-out composite, never a scoring input
- Internal Database Migration Tasks ; OpenAI: 63.9% · Contributes. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Internal Design Tasks ; OpenAI: 50% · Contributes. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Internal Data Science Tasks ; OpenAI: 40.9% · Contributes. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- OpenScore String Quartets (1 - OMR-NED): 0.84 · Not used. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- MedChemBench (internal): 49.3% · Not used. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- MRCR v2 (8 needles, 512K-1M): 96.3% · Not used. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- ExploitBench (Jun-Aug 2026, 300-turn limit): 39% · Not used. Source · Developer-reported · GPT-6 Astra; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Mystery Game Puzzles v1.0.4 ; Epoch: 84% · Contributes. Source · Evaluator-reported · gpt-6-astra_max; task version 1.0.4
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- WANDR ; Perplexity: 0.682 · Not used.
Exact reported metric and evaluation configuration unresolved.
- FrontierMath Erdos (68 conjectures) ; Epoch: 2.941176 · Not used. Source · Evaluator-reported · gpt-6-astra_max; task version
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- Box Complex Work Eval (overall): 77% · Not used.
Not admitted after the September 6 model review: the collected 77 for GPT-6 Astra on Box Complex Work Eval (overall) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCyber (Sep 2026) ; Irregular: 86 · Not used.
Not admitted after the September 6 model review: the collected 86 for GPT-6 Astra on FrontierCyber (Sep 2026) ; Irregular still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CyScenarioBench (Sep 2026) ; Irregular: 59% · Not used.
Not admitted after the September 6 model review: the collected 59 for GPT-6 Astra on CyScenarioBench (Sep 2026) ; Irregular still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- BioMysteryBench v1 ; Vals: 79.26% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Tax Agent Bench v1 ; Vals: 63.34% · Contributes. Source · Evaluator-reported · openai_gpt-6-astra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- RuneBench (30m, mean ln(1 + XP/min)): 7.258225 · Contributes. Source · Archived source review · gpt6astra-high
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 2714 · Contributes. Source · Archived source review · GPT-6 Astra (Max)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Roboflow Vision Evals (six-task mean, high tier): 88.533333% · Contributes. Source · Archived source review · Require a complete six-task high-tier panel, the highest published benchmark tier. Retain high-tier scores even when lower than low-tier. Do not substitute the best of three repetitions: use the published mean. Each panel enters the fit once; do not also fit its components.
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1797 · Contributes. Source · Archived source review · gpt-6-astra-max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- AA-AnalystAgent: 51.25% · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); Artificial Analysis published evaluation harness
Exact checkpoint, metric and published aggregate matched to the evaluator source. Date is retrieval date; first publication date is not inferred.
- AA-LCR v1.1: 80.666667% · Not used. Source · Evaluator-reported · GPT-6 Astra (max)
AA v4.3 evaluation, kept separate from prior benchmark versions and other agent harnesses.
- Agent Arena overall IPS: 0.125451 · Not used. Source · Evaluator-reported · GPT-6 Astra (Max); Agent Arena 2026-09-08
Native inverse-propensity-score metric. Retained as a separate uncalibrated benchmark, not converted to an accuracy percentage or Text Arena Elo.
- Astra-26 LiveBrowseComp ; native web: 42.307692% · Not used. Source · Evaluator-reported · GPT-6 Astra medium; Codex CLI 0.153.4; built-in web
Native web reference configuration only. Exa, Keenable, Parallel, and Valyu are different tool stacks, not interchangeable reasoning modes. One run per task; not the existing BrowseComp benchmark.
- AutomationBench-AA: 68.491748% · Not used. Source · Evaluator-reported · GPT-6 Astra (max)
AA v4.3 evaluation, kept separate from prior benchmark versions and other agent harnesses.
- Creative Writing v3: 2163.9 · Contributes. Source · Archived source review · gpt-6-astra; published evaluation, effort unspecified
Star denotes community-submitted evaluation hosted on the official benchmark board; not independently rerun by the benchmark author.
- GDP.pdf ; AA harness: 32.2% · Not used. Source · Evaluator-reported · GPT-6 Astra (xhigh)
AA v4.3 evaluation, kept separate from prior benchmark versions and other agent harnesses.
- LiveBench: 82.160643% · Contributes. Source · Evaluator-reported · gpt-6-astra-max; LiveBench 2026-06-25 task set
Derived exactly using the existing category mapping and evaluator-published task scores; not a new benchmark version.
- MathArena ; ArXivLean Jun 2026: 63.83% · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); MathArena June 2026 competition harness
Exact checkpoint, metric and published aggregate matched to the evaluator source. Date is retrieval date; first publication date is not inferred.
- MathArena ; ArXivMath Jun 2026: 94.44% · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); MathArena named monthly competition
Exact published two-decimal accuracy from the direct model comparison. Month is kept explicit; overall values are not mixed with month-level rows.
- MathArena ; ArXivMath May 2026: 95% · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); MathArena named monthly competition
Exact published two-decimal accuracy from the direct model comparison. Month is kept explicit; overall values are not mixed with month-level rows.
- MathArena ; BrokenArXiv Jun 2026: 99.07% · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); MathArena named monthly competition
Exact published two-decimal accuracy from the direct model comparison. Month is kept explicit; overall values are not mixed with month-level rows.
- MLCR-AA (Medical Long Context Reasoning): 35% · Contributes. Source · Evaluator-reported · GPT-6 Astra (max); Artificial Analysis published evaluation harness
Exact checkpoint, metric and published aggregate matched to the evaluator source. Date is retrieval date; first publication date is not inferred.
- Terminal-Bench 4.0 ; AA harness: 59.59596% · Not used. Source · Evaluator-reported · GPT-6 Astra (xhigh)
AA v4.3 evaluation, kept separate from prior benchmark versions and other agent harnesses.
Fable 5.1 · 40 used / 74 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
36 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.4 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1763.64 · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Source still labels this model Default Fallback. Retained for completeness, excluded from single-model scoring.
- Terminal-Bench 2.1: 91.4% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- DeepSWE v1.1: 67.4% · Contributes. Source · Developer-reported · Fable 5.1 final snapshot; adaptive max; five trials
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LMArena (Chatbot Arena) ; Text: 1504.213301 · Contributes. Source · Evaluator-reported · claude-fable-5.1-max; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 66 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 93.74% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- AA-LCR: 80% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 58.88% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 76.67% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 63.078704% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Source still labels this model Default Fallback. Retained for completeness, excluded from single-model scoring.
- Humanity's Last Exam (no tools): 59.128823% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Source still labels this model Default Fallback. Retained for completeness, excluded from single-model scoring.
- Legal Research Bench ; Vals: 55.29% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 6.67% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1694 · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- MedScribe ; Vals: 91.29% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 88.51% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 75.96% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Public Benefits Bench ; Vals: 74.9% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 53.51% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Omniscience: 43.45 · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Source still labels this model Default Fallback. Retained for completeness, excluded from single-model scoring.
- MMLU Pro ; Vals: 92.38% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 83.414155% · Contributes. Source · Evaluator-reported · claude-fable-5-1-max-effort; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- MortgageTax ; Vals: 70.79% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 48.53% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Toolathlon-Verified: 77.8% · Not used.
Not admitted after the September 6 model review: the collected 77.8 for Fable 5.1 on Toolathlon-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- LiveCodeBench: 90.52% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Humanity's Last Exam w/ tools: 65% · Contributes. Source · Developer-reported · Fable 5.1; OpenAI comparison table; maximum across reasoning modes
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 82.688% · Contributes. Source · Archived source review · anthropic/claude-fable-5-1 · compute=max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- FrontierSWE: 57% · Not used.
Not admitted after the September 6 model review: the collected 57 for Fable 5.1 on FrontierSWE still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierMath v2 Tier 4: 87.804878% · Contributes. Source · Evaluator-reported · claude-fable-5-1_max; task version 2.0.0
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- CursorBench v3.2: 73.4% · Contributes. Source · Evaluator-reported · Fable 5.1 Max; Cursor harness
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- APEX-Agents Mean Criteria Passed: 77.1% · Contributes. Source · Evaluator-reported · claude-fable-5.1; max; loop_truncated_tools_agent
Same 240-task, 31-world APEX benchmark and mean-score metric. Current published repeated-run aggregate (959 samples); pass@1 is not substituted.
- SWE-bench Pro: 81.2% · Not used.
Not admitted after the September 6 model review: the collected 81.2 for Fable 5.1 on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- APEX-SWE (Pass@1, Terminus-2): 63.6% · Contributes. Source · Evaluator-reported · claude-fable-5.1; max; terminus-2
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MMMU-Pro: 90.64% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- ARC-AGI-2: 90% · Contributes. Source · Evaluator-reported · Fable 5.1; XHigh; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- OSWorld 2.0 ; Anthropic H2H: 77.9% · Not used.
Not admitted after the September 6 model review: the collected 77.9 for Fable 5.1 on OSWorld 2.0 ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Extended: 63.6% · Contributes. Source · Evaluator-reported · Claude Fable 5.1; medium; claude-code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MathArena ; ArXivMath Jun 2026: 90.97% · Contributes. Source · Evaluator-reported · Fable 5.1 (max); MathArena named monthly competition
Exact published two-decimal accuracy from the direct model comparison. Month is kept explicit; overall values are not mixed with month-level rows.
- OfficeQA Pro †: 69% · Not used.
Not admitted after the September 6 model review: the collected 69 for Fable 5.1 on OfficeQA Pro † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Main ; Anthropic H2H: 50.9% · Not used.
Not admitted after the September 6 model review: the collected 50.9 for Fable 5.1 on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AutomationBench ; Anthropic H2H: 31.4% · Not used.
Not admitted after the September 6 model review: the collected 31.4 for Fable 5.1 on AutomationBench ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Harvey Legal Agent Benchmark ; held-out: 16.7% · Not used.
Not admitted after the September 6 model review: the collected 16.7 for Fable 5.1 on Harvey Legal Agent Benchmark ; held-out still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 47.2% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- CritPt: 29.7% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- Terminal-Bench 4.0: 57.9% · Contributes. Source · Evaluator-reported · Fable 5.1; max; Claude Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 100% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 90.26% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5-1; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 2050 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- ARC-AGI-1: 97.5% · Contributes. Source · Evaluator-reported · Fable 5.1; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- BenchCAD (Vision2Code subset): 84.3% · Not used. Source · Developer-reported · Fable 5.1; max; with tools; 1000-file subset; three disclosed reference modifications
Matched 1000-file Vision2Code subset and disclosed modified grading. Full-set OpenAI results cannot be silently compared as this subset.
- Terminal-Bench Science 0.1: 52.6% · Contributes. Source · Developer-reported · Fable 5.1; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Harvey LAB-AA: 93.02% · Not used.
Not admitted after the September 6 model review: the collected 93.02 for Fable 5.1 on Harvey LAB-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vending-Bench 2: 5421.556667 · Contributes. Source · Evaluator-reported · Claude Fable 5.1; Andon Labs Vending-Bench 2
Compare reasoning configurations using their mean over runs. Geometric mean, individual best runs, Fireworks Kimi endpoint, Vending Arena, and April DeepSeek Pro are not substituted.
- Internal Database Migration Tasks ; OpenAI: 57.8% · Contributes. Source · Developer-reported · Fable 5.1; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Mystery Game Puzzles v1.0.4 ; Epoch: 58% · Contributes. Source · Evaluator-reported · claude-fable-5-1_max; task version 1.0.4
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- FrontierMath Erdos (68 conjectures) ; Epoch: 0 · Not used. Source · Evaluator-reported · claude-fable-5-1_max; task version
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- RuneBench (30m, mean ln(1 + XP/min)): 6.143277 · Contributes. Source · Archived source review · fable51-xhigh
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 2245 · Contributes. Source · Archived source review · Claude Fable 5.1
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Roboflow Vision Evals (six-task mean, high tier): 82.516667% · Contributes. Source · Archived source review · Require a complete six-task high-tier panel, the highest published benchmark tier. Retain high-tier scores even when lower than low-tier. Do not substitute the best of three repetitions: use the published mean. Each panel enters the fit once; do not also fit its components.
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1762 · Contributes. Source · Archived source review · claude-fable-5.1-max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- AA-AnalystAgent: 57.5% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Source still labels this model Default Fallback. Retained for completeness, excluded from single-model scoring.
- AA-LCR v1.1: 85.333333% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Distinct AA protocol/version; not pooled with other harnesses. Fable Default Fallback scores are retained but not attributed to a single checkpoint.
- Agent Arena overall IPS: 0.145069 · Not used. Source · Evaluator-reported · Fable 5.1 (Max); Agent Arena 2026-09-08
Native inverse-propensity-score metric. Retained as a separate uncalibrated benchmark, not converted to an accuracy percentage or Text Arena Elo.
- AutomationBench-AA: 59.375917% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Distinct AA protocol/version; not pooled with other harnesses. Fable Default Fallback scores are retained but not attributed to a single checkpoint.
- Code Migration †: 54.607% · Not used. Source · Evaluator-reported · Claude Fable 5.1; max; Vals published harness
Newly located exact Vals overall result. Retained as unresolved because Fable model page discloses fallback routing and benchmark-level impact has not been established.
- Creative Writing v3: 2152.7 · Contributes. Source · Archived source review · claude-fable-5-1; published evaluation, effort unspecified
Star denotes community-submitted evaluation hosted on the official benchmark board; not independently rerun by the benchmark author.
- GDP.pdf ; AA harness: 28% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)
Distinct AA protocol/version; not pooled with other harnesses. Fable Default Fallback scores are retained but not attributed to a single checkpoint.
- MathArena ; ArXivMath May 2026: 87.5% · Contributes. Source · Evaluator-reported · Fable 5.1 (max); MathArena named monthly competition
Exact published two-decimal accuracy from the direct model comparison. Month is kept explicit; overall values are not mixed with month-level rows.
- MathArena ; BrokenArXiv Jun 2026: 84.72% · Contributes. Source · Evaluator-reported · Fable 5.1 (max); MathArena named monthly competition
Exact published two-decimal accuracy from the direct model comparison. Month is kept explicit; overall values are not mixed with month-level rows.
- MLCR-AA (Medical Long Context Reasoning): 71.111111% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Source still labels this model Default Fallback. Retained for completeness, excluded from single-model scoring.
- SkillsBench †: 61.552% · Not used. Source · Evaluator-reported · Claude Fable 5.1; max; Vals published harness
Newly located exact Vals overall result. Retained as unresolved because Fable model page discloses fallback routing and benchmark-level impact has not been established.
- Tax Agent Bench v1 ; Vals: 77.642% · Not used. Source · Evaluator-reported · Claude Fable 5.1; max; Vals published harness
Newly located exact Vals overall result. Retained as unresolved because Fable model page discloses fallback routing and benchmark-level impact has not been established.
- Terminal-Bench 4.0 ; AA harness: 55.050505% · Not used. Source · Evaluator-reported · Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
Distinct AA protocol/version; not pooled with other harnesses. Fable Default Fallback scores are retained but not attributed to a single checkpoint.
Opus 5 · 69 used / 94 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
59 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.1 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1738.08 · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 89.138577% · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 73.648649% · Contributes. Source · Evaluator-reported · claude-opus-5; max; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1493.32766 · Contributes. Source · Evaluator-reported · claude-opus-5-high; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 63 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 93.737374% · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, Xhigh Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 75.7% · Not used. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, Max Effort)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 58.63% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 73.56% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 56.365741% · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 54.865616% · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 55.29% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 6.67% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1647.02 · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 90.98% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 97% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 86.97% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 75.14% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- KingBench 3: 77.5% · Not used.
Not admitted after the September 6 model review: the collected 77.5 for Opus 5 on KingBench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Index: 74.8% · Not used. Source · Evaluator-reported · claude-opus-5; max effort; Vals AI index methodology
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 76.93% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 63.57% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AutomationBench v1.0.6: 50.3% · Not used.
Not admitted after the September 6 model review: the collected 50.3 for Opus 5 on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 37.066667 · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 91.59% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 80.084976% · Contributes. Source · Evaluator-reported · claude-opus-5-max-effort; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- CorpFin v2 ; Vals: 73.2% · Contributes. Source · Archived source review · claude-opus-5; max effort; Vals AI CorpFin v2 evaluation with Sonnet 4.5 judge
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MortgageTax ; Vals: 72.06% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 49.43% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 57.47% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Toolathlon-Verified: 80.6% · Not used.
Not admitted after the September 6 model review: the collected 80.6 for Opus 5 on Toolathlon-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Agents' Last Exam: 31.6% · Contributes. Source · Evaluator-reported · Opus 5; High; Claude Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Terminal-Bench v3.0: 42.7% · Contributes. Source · Evaluator-reported · Claude Opus 5 (max); mini-SWE-agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveCodeBench: 89.03% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Humanity's Last Exam w/ tools: 63.6% · Contributes. Source · Developer-reported · Opus 5; OpenAI comparison table; maximum across reasoning modes
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Design Arena (Elo): 1407 · Contributes. Source · Archived source review · claude-opus-5; max effort; Design Arena 2026-08-11 rating snapshot
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Vals Multimodal Index: 73.9% · Not used. Source · Evaluator-reported · claude-opus-5; max effort; Vals AI multimodal index methodology
Held-out composite, never a scoring input
- SWE-Marathon v1.1: 50% · Not used.
Not admitted after the September 6 model review: the collected 50 for Opus 5 on SWE-Marathon v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 82.273% · Contributes. Source · Archived source review · anthropic/claude-opus-5 · compute=max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- SkillsBench †: 60.44% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierMath v2 Tier 4: 73.170732% · Contributes. Source · Evaluator-reported · claude-opus-5_max; task version 2.0.0
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- CursorBench v3.2: 70% · Contributes. Source · Evaluator-reported · Opus 5 Max; Cursor harness
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- APEX-Agents Mean Criteria Passed: 60.6% · Contributes. Source · Evaluator-reported · claude-opus-5; max; loop_truncated_tools_agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Pro: 79.2% · Not used.
Not admitted after the September 6 model review: the collected 79.2 for Opus 5 on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CharXiv (with CI / RQ): 89.3% · Not used.
Not admitted after the September 6 model review: the collected 89.3 for Opus 5 on CharXiv (with CI / RQ) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- APEX-SWE (Pass@1, Terminus-2): 63.7% · Contributes. Source · Evaluator-reported · claude-opus-5; max; terminus-2
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- BrowseComp: 90.8% · Contributes. Source · Developer-reported · Opus 5; OpenAI comparison table; maximum across reasoning modes
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MMMU-Pro: 84.739884% · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ARC-AGI-2: 90.4% · Contributes. Source · Evaluator-reported · Opus 5; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- OSWorld 2.0 ; Anthropic H2H: 70.6% · Contributes. Source · Archived source review · claude-opus-5; max effort; Anthropic Opus 5 published head-to-head setup
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- ExploitBench: 70% · Not used.
Inherited comparator metric/scaffold equivalence to Astra Cap Percent is not established; the source flags possible contamination. Excluded pending protocol audit.
- PostTrainBench: 35.04% · Not used.
Not admitted after the September 6 model review: the collected 35.04 for Opus 5 on PostTrainBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SimpleQA: 56.7% · Not used.
Not admitted after the September 6 model review: the collected 56.7 for Opus 5 on SimpleQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Extended: 63.63% · Contributes. Source · Evaluator-reported · Claude Opus 5; medium; claude-code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- JobBench: 65.7% · Not used.
Not admitted after the September 6 model review: the collected 65.7 for Opus 5 on JobBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- OSWorld-Verified: 90.69% · Not used.
Not admitted after the September 6 model review: the collected 90.69 for Opus 5 on OSWorld-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Creative Writing v3: 2116.1 · Contributes. Source · Evaluator-reported · claude-opus-5; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- MathArena ; ArXivMath Jun 2026: 80.95% · Contributes. Source · Archived source review · claude-opus-5; max effort; MathArena June 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MathArena ; BrokenArXiv Jun 2026: 90.74% · Contributes. Source · Archived source review · claude-opus-5; max effort; MathArena June 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- AIIQ Composite IQ: 134 · Not used. Source · Evaluator-reported · claude-opus-5; max effort; AIIQ composite snapshot
Held-out composite, never a scoring input
- MCP Atlas †: 85.8% · Not used.
Not admitted after the September 6 model review: the collected 85.8 for Opus 5 on MCP Atlas † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- OfficeQA Pro †: 66.9% · Not used.
Not admitted after the September 6 model review: the collected 66.9 for Opus 5 on OfficeQA Pro † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EQ-Bench 4: 1385 · Contributes. Source · Evaluator-reported · claude-opus-5; evaluator published configuration, effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MathArena ; ArXivLean Jun 2026: 31.25% · Contributes. Source · Archived source review · claude-opus-5; max effort; MathArena June 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Frontier-Bench v0.1 ; Anthropic H2H: 43.3% · Contributes. Source · Archived source review · claude-opus-5; max effort; Anthropic Opus 5 published head-to-head setup
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- FrontierCode v1.1 Main ; Anthropic H2H: 53.4% · Not used.
Not admitted after the September 6 model review: the collected 53.4 for Opus 5 on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AutomationBench ; Anthropic H2H: 26% · Contributes. Source · Archived source review · claude-opus-5; max effort; Anthropic Opus 5 published head-to-head setup
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Harvey Legal Agent Benchmark ; held-out: 11.7% · Not used.
Not admitted after the September 6 model review: the collected 11.7 for Opus 5 on Harvey Legal Agent Benchmark ; held-out still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 44.742268% · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, High Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 29.142857% · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 4.0: 51.8% · Contributes. Source · Evaluator-reported · Opus 5; max; Claude Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- GDP.pdf: 24% · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive/Max); Surge original evaluation
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 99% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 88.4% · Contributes. Source · Evaluator-reported · anthropic_claude-opus-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 2106 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- OSWorld 2.0: 31.43% · Not used.
Not admitted after the September 6 model review: the collected 31.43 for Opus 5 on OSWorld 2.0 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ARC-AGI-1: 97.5% · Contributes. Source · Evaluator-reported · Opus 5; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- ARC-AGI-3 (standard harness): 30.16% · Contributes. Source · Evaluator-reported · Opus 5; High; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- Terminal-Bench Science 0.1: 30% · Contributes. Source · Developer-reported · Opus 5; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- SRE-Bench (pass@1): 12.5% · Contributes. Source · Developer-reported · Opus 5; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- OSWorld 2.0 ; OpenAI H2H (offline subset): 70.2% · Contributes. Source · Developer-reported · Opus 5; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Agents' Last Exam ; OpenAI H2H: 55.5% · Contributes. Source · Developer-reported · Opus 5; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- HealthBench Professional (length-adjusted): 56.4% · Contributes. Source · Developer-reported · Opus 5; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Harvey LAB-AA: 93.46% · Not used.
Not admitted after the September 6 model review: the collected 93.46 for Opus 5 on Harvey LAB-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EnterpriseOps-Gym-AA: 47.48% · Not used.
Not admitted after the September 6 model review: the collected 47.48 for Opus 5 on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 59.444444% · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, High Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-AnalystAgent: 53.75% · Contributes. Source · Evaluator-reported · Claude Opus 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSearchQA: 95% · Not used.
Not admitted after the September 6 model review: the collected 95 for Opus 5 on DeepSearchQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vending-Bench 2: 11181.873333 · Contributes. Source · Evaluator-reported · Claude Opus 5; Andon Labs Vending-Bench 2
Compare reasoning configurations using their mean over runs. Geometric mean, individual best runs, Fireworks Kimi endpoint, Vending Arena, and April DeepSeek Pro are not substituted.
- ZeroBench: 26% · Contributes. Source · Evaluator-reported · Claude Opus 5 (max); official main questions
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Mystery Game Puzzles v1.0.4 ; Epoch: 59% · Contributes. Source · Evaluator-reported · claude-opus-5_max; task version 1.0.4
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- RuneBench (30m, mean ln(1 + XP/min)): 5.705729 · Contributes. Source · Archived source review · opus5-xhigh
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 2315 · Contributes. Source · Archived source review · Claude Opus 5 (Max)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1688 · Contributes. Source · Archived source review · claude-opus-5-max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
Fable 5 · 57 used / 114 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
48 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.2 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1741 · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- Terminal-Bench 2.1: 84.6% · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- DeepSWE v1.1: 69.911504% · Contributes. Source · Evaluator-reported · claude-fable-5; xhigh; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1507.164172 · Contributes. Source · Evaluator-reported · claude-fable-5; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 62 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 92.6% · Contributes. Source · Archived source review · claude-fable-5; max effort; No tools
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- AA-LCR: 70% · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 56.31% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 73.67% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 60% · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- Humanity's Last Exam (no tools): 53.3% · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- Legal Research Bench ; Vals: 49.52% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 11.25% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1574 · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- MedScribe ; Vals: 88.52% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 95% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 88.56% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 76.94% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- KingBench 3: 82.5% · Not used.
Not admitted after the September 6 model review: the collected 82.5 for Fable 5 on KingBench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Index: 75.1% · Not used. Source · Evaluator-reported · claude-fable-5; max effort; Vals AI index methodology
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 70.43% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 56.07% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AutomationBench v1.0.6: 46.2% · Not used.
Not admitted after the September 6 model review: the collected 46.2 for Fable 5 on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 40.15 · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- MMLU Pro ; Vals: 91.5% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 82.970798% · Contributes. Source · Evaluator-reported · claude-fable-5-max-effort; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- CorpFin v2 ; Vals: 71.8% · Contributes. Source · Archived source review · claude-fable-5; max effort; Vals AI CorpFin v2 evaluation with Sonnet 4.5 judge
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MortgageTax ; Vals: 68.92% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 51.89% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 55.06% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Toolathlon-Verified: 77.9% · Not used.
Not admitted after the September 6 model review: the collected 77.9 for Fable 5 on Toolathlon-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Agents' Last Exam: 25.7% · Contributes. Source · Evaluator-reported · Fable 5; XHigh; Claude Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Terminal-Bench v3.0: 34.1% · Contributes. Source · Evaluator-reported · Claude Fable 5 (max); Claude Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveCodeBench: 89.78% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CyberGym: 83.8% · Not used.
Not admitted after the September 6 model review: the collected 83.8 for Fable 5 on CyberGym still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Humanity's Last Exam w/ tools: 63.8% · Contributes. Source · Developer-reported · Fable 5; OpenAI comparison table; maximum across reasoning modes
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Design Arena (Elo): 1400 · Contributes. Source · Archived source review · claude-fable-5; max effort; Design Arena 2026-08-11 rating snapshot
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Vals Multimodal Index: 74.2% · Not used. Source · Evaluator-reported · claude-fable-5; max effort; Vals AI multimodal index methodology
Held-out composite, never a scoring input
- SWE-Marathon v1.1: 33.1% · Not used.
Not admitted after the September 6 model review: the collected 33.1 for Fable 5 on SWE-Marathon v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SkillsBench †: 70.9% · Not used.
Not admitted after the September 6 model review: the collected 70.9 for Fable 5 on SkillsBench † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierSWE: 86.6% · Not used.
Not admitted after the September 6 model review: the collected 86.6 for Fable 5 on FrontierSWE still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierMath v2 Tier 4: 87.804878% · Contributes. Source · Evaluator-reported · claude-fable-5_max; task version 2.0.0
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- CursorBench v3.2: 70.5% · Contributes. Source · Evaluator-reported · Fable 5 Max; Cursor harness
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- APEX-Agents Mean Criteria Passed: 59.2% · Contributes. Source · Evaluator-reported · claude-fable-5; max; loop_truncated_tools_agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Pro: 80% · Not used.
Not admitted after the September 6 model review: the collected 80 for Fable 5 on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CharXiv (with CI / RQ): 93.5% · Not used.
Not admitted after the September 6 model review: the collected 93.5 for Fable 5 on CharXiv (with CI / RQ) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- APEX-SWE (Pass@1, Terminus-2): 58.8% · Contributes. Source · Evaluator-reported · claude-fable-5; max; terminus-2
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- BrowseComp: 87.4% · Contributes. Source · Developer-reported · Fable 5; OpenAI comparison table; maximum across reasoning modes
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MMMU-Pro: 81.2% · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- ARC-AGI-2: 89.2% · Contributes. Source · Evaluator-reported · Fable 5; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- OSWorld 2.0 ; Anthropic H2H: 66.1% · Contributes. Source · Archived source review · claude-fable-5; max effort; Anthropic Opus 5 published head-to-head setup
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- ExploitBench: 78% · Not used.
Inherited comparator metric/scaffold equivalence to Astra Cap Percent is not established; the source flags possible contamination. Excluded pending protocol audit.
- PostTrainBench: 41.8% · Not used.
Not admitted after the September 6 model review: the collected 41.8 for Fable 5 on PostTrainBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SimpleQA: 68.3% · Not used.
Not admitted after the September 6 model review: the collected 68.3 for Fable 5 on SimpleQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Extended: 64.94% · Contributes. Source · Evaluator-reported · Claude Fable 5; xhigh; claude-code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- JobBench: 57.4% · Not used.
Not admitted after the September 6 model review: the collected 57.4 for Fable 5 on JobBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- BabyVision (with CI): 90.5% · Not used.
Not admitted after the September 6 model review: the collected 90.5 for Fable 5 on BabyVision (with CI) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- PerceptionBench: 57.2% · Not used.
Not admitted after the September 6 model review: the collected 57.2 for Fable 5 on PerceptionBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- LVBench (with Memory): 90.1% · Not used.
Not admitted after the September 6 model review: the collected 90.1 for Fable 5 on LVBench (with Memory) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- OSWorld-Verified: 85% · Not used.
Not admitted after the September 6 model review: the collected 85 for Fable 5 on OSWorld-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Creative Writing v3: 1933.2 · Contributes. Source · Evaluator-reported · claude-fable-5; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- MathArena ; ArXivMath Jun 2026: 83.67% · Contributes. Source · Archived source review · claude-fable-5; max effort; MathArena June 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MathArena ; BrokenArXiv Jun 2026: 47.84% · Contributes. Source · Archived source review · claude-fable-5; max effort; MathArena June 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- AIIQ Composite IQ: 134 · Not used. Source · Evaluator-reported · claude-fable-5; max effort; AIIQ composite snapshot
Held-out composite, never a scoring input
- MCP Atlas †: 84.7% · Not used.
Not admitted after the September 6 model review: the collected 84.7 for Fable 5 on MCP Atlas † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- OfficeQA Pro †: 69.9% · Not used.
Not admitted after the September 6 model review: the collected 69.9 for Fable 5 on OfficeQA Pro † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EQ-Bench 4: 1340.4 · Contributes. Source · Evaluator-reported · claude-fable-5; evaluator published configuration, effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CoWorkBench: 75.9% · Not used.
Not admitted after the September 6 model review: the collected 75.9 for Fable 5 on CoWorkBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ²-Bench Telecom: 98.5% · Not used.
Not admitted after the September 6 model review: the collected 98.5 for Fable 5 on τ²-Bench Telecom still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- PaperBench (Replication Score): 88.8% · Not used.
Not admitted after the September 6 model review: the collected 88.8 for Fable 5 on PaperBench (Replication Score) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- QwenReactBench (Elo): 1770 · Not used.
Not admitted after the September 6 model review: the collected 1770 for Fable 5 on QwenReactBench (Elo) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ERQA: 70% · Not used.
Not admitted after the September 6 model review: the collected 70 for Fable 5 on ERQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vision2Web (Avg. Frontend/Webpage/etc.): 70.5% · Not used.
Not admitted after the September 6 model review: the collected 70.5 for Fable 5 on Vision2Web (Avg. Frontend/Webpage/etc.) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MobileWorld: 85.5% · Not used.
Not admitted after the September 6 model review: the collected 85.5 for Fable 5 on MobileWorld still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- WebArena-Verified: 71.3% · Not used.
Not admitted after the September 6 model review: the collected 71.3 for Fable 5 on WebArena-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathArena ; ArXivLean Jun 2026: 14.58% · Contributes. Source · Archived source review · claude-fable-5; max effort; MathArena June 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Frontier-Bench v0.1 ; Anthropic H2H: 33.7% · Contributes. Source · Archived source review · claude-fable-5; max effort; Anthropic Opus 5 published head-to-head setup
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- FrontierCode v1.1 Main ; Anthropic H2H: 53.5% · Not used.
Not admitted after the September 6 model review: the collected 53.5 for Fable 5 on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AutomationBench ; Anthropic H2H: 17.4% · Contributes. Source · Archived source review · claude-fable-5; max effort; Anthropic Opus 5 published head-to-head setup
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Harvey Legal Agent Benchmark ; held-out: 13.3% · Not used.
Not admitted after the September 6 model review: the collected 13.3 for Fable 5 on Harvey Legal Agent Benchmark ; held-out still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathArena ; ArXivMath May 2026: 87.5% · Contributes. Source · Archived source review · claude-fable-5; max effort; MathArena May 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- EQ-Bench 3: 2049.7 · Not used. Source · Evaluator-reported · *claude-fable-5; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- τ³-Banking: 38.1% · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- CritPt: 28.6% · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- Terminal-Bench 4.0: 44.5% · Contributes. Source · Evaluator-reported · Fable 5; max; Claude Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- GDP.pdf: 29.8% · Contributes. Source · Evaluator-reported · Claude Fable 5 (Adaptive/Max); Surge original evaluation
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 95% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 90.35% · Contributes. Source · Evaluator-reported · anthropic_claude-fable-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 1972 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- ARC-AGI-1: 98.5% · Contributes. Source · Evaluator-reported · Fable 5; XHigh; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- Terminal-Bench Science 0.1: 21.4% · Contributes. Source · Developer-reported · Fable 5; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Agents' Last Exam ; OpenAI H2H: 48.7% · Contributes. Source · Developer-reported · Fable 5; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- HealthBench Professional (length-adjusted): 60.9% · Contributes. Source · Developer-reported · Fable 5; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Harvey LAB-AA: 93.56% · Not used.
Not admitted after the September 6 model review: the collected 93.56 for Fable 5 on Harvey LAB-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EnterpriseOps-Gym-AA: 51.12% · Not used.
Not admitted after the September 6 model review: the collected 51.12 for Fable 5 on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 64.44% · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- AA-AnalystAgent: 48.75% · Not used. Source · Evaluator-reported · Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
AA labels this configuration Default Fallback; per-evaluation fallback use is not established. Do not attribute a multi-model result to the named model alone.
- DeepSearchQA: 94.2% · Not used.
Not admitted after the September 6 model review: the collected 94.2 for Fable 5 on DeepSearchQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vending-Bench 2: 5680.258 · Contributes. Source · Evaluator-reported · Claude Fable 5 - High; Andon Labs Vending-Bench 2
Compare reasoning configurations using their mean over runs. Geometric mean, individual best runs, Fireworks Kimi endpoint, Vending Arena, and April DeepSeek Pro are not substituted.
- IFBench: 63.5% · Not used.
Not admitted after the September 6 model review: the collected 63.5 for Fable 5 on IFBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AndroidWorld: 88.8% · Not used.
Not admitted after the September 6 model review: the collected 88.8 for Fable 5 on AndroidWorld still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CWE-bench: 47.8% · Not used.
Not admitted after the September 6 model review: the collected 47.8 for Fable 5 on CWE-bench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ZeroBench: 24% · Contributes. Source · Evaluator-reported · Claude Fable 5 (max); official main questions
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SimpleVQA: 73.4% · Not used.
Not admitted after the September 6 model review: the collected 73.4 for Fable 5 on SimpleVQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathVision (with CI): 98.6% · Not used.
Not admitted after the September 6 model review: the collected 98.6 for Fable 5 on MathVision (with CI) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- OmniDocBench 1.5: 89.5% · Not used.
Not admitted after the September 6 model review: the collected 89.5 for Fable 5 on OmniDocBench 1.5 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Internal Database Migration Tasks ; OpenAI: 50.3% · Contributes. Source · Developer-reported · Fable 5; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Internal Design Tasks ; OpenAI: 35.8% · Contributes. Source · Developer-reported · Fable 5; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Internal Data Science Tasks ; OpenAI: 34.7% · Contributes. Source · Developer-reported · Fable 5; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Mystery Game Puzzles v1.0.4 ; Epoch: 52% · Contributes. Source · Evaluator-reported · claude-fable-5_max; task version 1.0.4
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- FrontierMath Erdos (68 conjectures) ; Epoch: 0 · Not used. Source · Evaluator-reported · claude-fable-5_max; task version
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- RuneBench (30m, mean ln(1 + XP/min)): 5.67775 · Contributes. Source · Archived source review · fable-5-xhigh
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 2149 · Contributes. Source · Archived source review · Claude Fable 5 (Max)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1629 · Contributes. Source · Archived source review · claude-fable-5
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
Muse Spark 1.3 · 13 used / 23 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
13 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±18.2 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1719.65 · Contributes. Source · Evaluator-reported · Muse Spark 1.3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 85.76779% · Contributes. Source · Evaluator-reported · Muse Spark 1.3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 75.4% · Not used.
Not admitted after the September 6 model review: the collected 75.4 for Muse Spark 1.3 on DeepSWE v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA Intelligence Index v4.1.1 (score): 61 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 94.141414% · Contributes. Source · Evaluator-reported · Muse Spark 1.3 (xhigh); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 79% · Not used. Source · Evaluator-reported · Muse Spark 1.3 (xhigh)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- SciCode: 59.722222% · Contributes. Source · Evaluator-reported · Muse Spark 1.3 (xhigh); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 49.073216% · Contributes. Source · Evaluator-reported · Muse Spark 1.3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AutomationBench v1.0.6: 43.8% · Not used.
Not admitted after the September 6 model review: the collected 43.8 for Muse Spark 1.3 on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 24.933333 · Contributes. Source · Evaluator-reported · Muse Spark 1.3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- LiveBench: 81.59% · Not used.
Not admitted after the September 6 model review: the collected 81.59 for Muse Spark 1.3 on LiveBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MMMU-Pro: 82.023121% · Contributes. Source · Evaluator-reported · Muse Spark 1.3 (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- OSWorld 2.0 ; Anthropic H2H: 66.9% · Not used.
Not admitted after the September 6 model review: the collected 66.9 for Muse Spark 1.3 on OSWorld 2.0 ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- JobBench: 61.2% · Not used.
Not admitted after the September 6 model review: the collected 61.2 for Muse Spark 1.3 on JobBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 52.371134% · Contributes. Source · Evaluator-reported · Muse Spark 1.3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 26% · Contributes. Source · Evaluator-reported · Muse Spark 1.3 (xhigh); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MineBench: 1836 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- OSWorld 2.0: 32% · Not used.
Not admitted after the September 6 model review: the collected 32 for Muse Spark 1.3 on OSWorld 2.0 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Harvey LAB-AA: 95.48% · Not used.
Not admitted after the September 6 model review: the collected 95.48 for Muse Spark 1.3 on Harvey LAB-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Tax Agent Bench v1 ; Vals: 71.93% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_3; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- RuneBench (30m, mean ln(1 + XP/min)): 4.58575 · Contributes. Source · Archived source review · muse13
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Roboflow Vision Evals (six-task mean, high tier): 77.733333% · Contributes. Source · Archived source review · Require a complete six-task high-tier panel, the highest published benchmark tier. Retain high-tier scores even when lower than low-tier. Do not substitute the best of three repetitions: use the published mean. Each panel enters the fit once; do not also fit its components.
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1622 · Contributes. Source · Archived source review · muse-spark-1.3 (xHigh)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
DeepSeek V4.1-Flash · 10 used / 20 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
7 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±23.6 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GPQA Diamond: 90.9% · Contributes. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim.
- Humanity's Last Exam (no tools): 36.8% · Contributes. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim. Full HLE 36.8; the separately reported 39.1 is text-only and is not substituted.
- HLE text-only ; DeepSeek V4.1 release: 39.1% · Not used. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim. Retained as a separate protocol; does not contribute to the current frozen score.
- Codeforces Rating ; DeepSeek V4.1 release: 3471 · Not used. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim. Retained as a separate protocol; does not contribute to the current frozen score.
- MathArena Apex: 65.6% · Not used. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim. Retained as a separate protocol; does not contribute to the current frozen score.
- Terminal-Bench 2.1: 90.6% · Contributes. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim.
- Terminal-Bench v3.0: 30% · Contributes. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim.
- Terminal-Bench 4.0: 31.2% · Contributes. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim.
- DeepSWE v1.1: 74.2% · Contributes. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim.
- ProgramBench ; DeepSeek V4.1 release: 20.3% · Not used. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim. Retained as a separate protocol; does not contribute to the current frozen score.
- NL2Repo: 65.4% · Contributes. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim.
- CyberGym: 88.1% · Contributes. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim.
- SEC-Bench Pro ; DeepSeek V4.1 release: 62.8% · Not used. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim. Retained as a separate protocol; does not contribute to the current frozen score.
- ExploitGym: 15.3% · Not used. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim.
- Humanity's Last Exam w/ tools: 63.9% · Contributes. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim.
- Automation-Bench ; DeepSeek V4.1 release: 54.8% · Not used. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim. Retained as a separate protocol; does not contribute to the current frozen score.
- Agents' Last Exam: 31.8% · Contributes. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim.
- Chartography (w/tools): 78.9% · Not used. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim. Retained as a separate protocol; does not contribute to the current frozen score.
- BabyVision (w/tools): 89.6% · Not used. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim. Retained as a separate protocol; does not contribute to the current frozen score.
- ZeroBench-main (w/tools): 49% · Not used. Source · Evaluator-reported · DeepSeek V4.1-Flash, September 10 release; sole published setting (effort undisclosed)
Developer-reported release-table aggregate, corroborated against the supplied screenshot. Effort, harness, retry budget and fallback accounting are undisclosed; no maximum-effort claim. Retained as a separate protocol; does not contribute to the current frozen score.
GPT-5.6 Sol · 88 used / 135 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
75 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±0.9 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1627.48 · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 89.513109% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (xhigh); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 72.666667% · Contributes. Source · Evaluator-reported · gpt-5-6-sol; max; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1483.151082 · Contributes. Source · Evaluator-reported · gpt-5.6-sol-xhigh; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 61 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 94.141414% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 73.7% · Not used. Source · Evaluator-reported · GPT-5.6 Sol (max)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 53.76% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 72.34% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 57.75463% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (high); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 49.490269% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 48.08% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 2.5% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1478.66 · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 85.23% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 96.2% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 86.97% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 74.78% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- KingBench 3: 71.25% · Not used.
Not admitted after the September 6 model review: the collected 71.25 for GPT-5.6 Sol on KingBench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Index: 73.1% · Not used. Source · Evaluator-reported · gpt-5.6-sol; max effort; Vals AI index methodology
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 66.51% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 43.97% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AutomationBench v1.0.6: 45.8% · Not used.
Not admitted after the September 6 model review: the collected 45.8 for GPT-5.6 Sol on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 21.966667 · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 89.1% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 81.053548% · Contributes. Source · Evaluator-reported · gpt-5.6-sol-max; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- CorpFin v2 ; Vals: 64.4% · Contributes. Source · Archived source review · gpt-5.6-sol; max effort; Vals AI CorpFin v2 evaluation with Sonnet 4.5 judge
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MortgageTax ; Vals: 67.29% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 52.56% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 52.92% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Toolathlon-Verified: 74.9% · Not used.
Not admitted after the September 6 model review: the collected 74.9 for GPT-5.6 Sol on Toolathlon-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Agents' Last Exam: 30.6% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol; XHigh; Codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Terminal-Bench v3.0: 34.6% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveCodeBench: 82.6% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CyberGym: 83.6% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol; Z.ai GLM-5.3 model-card comparison table; GLM max, other comparator efforts not fully disclosed
CyberGym release-table percentage. Z.ai documents single-run pass@1 on 1,507 tasks, unlimited task timeout, Claude Code 2.1.207 and restricted network for GLM-5.3. Do not transfer those GLM-specific settings to every comparator. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- Humanity's Last Exam w/ tools: 64.5% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol; Z.ai GLM-5.3 model-card comparison table; GLM max, other comparator efforts not fully disclosed
HLE with tools release-table percentage. Z.ai documents temperature=1, top_p=.95, 163,840 output tokens, 300K context and GPT-5.6 Luna medium judge for its evaluation. Comparator-specific tools and judging settings are not fully disclosed. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- Design Arena (Elo): 1379 · Contributes. Source · Archived source review · gpt-5.6-sol; max effort; Design Arena 2026-08-11 rating snapshot
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Vals Multimodal Index: 72.2% · Not used. Source · Evaluator-reported · gpt-5.6-sol; max effort; Vals AI multimodal index methodology
Held-out composite, never a scoring input
- SWE-Marathon v1.1: 42.5% · Not used.
Not admitted after the September 6 model review: the collected 42.5 for GPT-5.6 Sol on SWE-Marathon v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 77.644% · Contributes. Source · Archived source review · openai/gpt-5.6-sol · reasoning_effort=max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- SkillsBench †: 54.1% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierSWE: 71.3% · Contributes. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierMath v2 Tier 4: 82.926829% · Contributes. Source · Evaluator-reported · gpt-5.6-sol_max; task version 2.0.0
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- CursorBench v3.2: 67.2% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol Max; Cursor harness
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- APEX-Agents Mean Criteria Passed: 56.7% · Contributes. Source · Evaluator-reported · gpt-5-6-sol-max; max; loop_truncated_tools_agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Pro: 64.6% · Not used.
Not admitted after the September 6 model review: the collected 64.6 for GPT-5.6 Sol on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CharXiv (with CI / RQ): 89.1% · Contributes. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; with Python tools
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- APEX-SWE (Pass@1, Terminus-2): 45.8% · Contributes. Source · Evaluator-reported · gpt-5-6-sol-xhigh; xhigh; terminus-2
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- BrowseComp: 90.4% · Contributes. Source · Developer-reported · GPT-5.6 Sol; OpenAI comparison table; maximum across reasoning modes
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MMMU-Pro: 83.410405% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ARC-AGI-2: 92.5% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- OSWorld 2.0 ; Anthropic H2H: 62.6% · Contributes. Source · Archived source review · gpt-5.6-sol; max effort; Anthropic Opus 5 published head-to-head setup
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- ExploitBench: 76.5% · Not used.
Inherited comparator metric/scaffold equivalence to Astra Cap Percent is not established; the source flags possible contamination. Excluded pending protocol audit.
- PostTrainBench: 34.6% · Contributes. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SimpleQA: 71.6% · Not used.
Not admitted after the September 6 model review: the collected 71.6 for GPT-5.6 Sol on SimpleQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Extended: 60.55% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol; max; codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- JobBench: 45.4% · Contributes. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- BabyVision (with CI): 88.9% · Contributes. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; with Python tools
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- PerceptionBench: 59.7% · Contributes. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LVBench (with Memory): 84.2% · Not used.
Not admitted after the September 6 model review: the collected 84.2 for GPT-5.6 Sol on LVBench (with Memory) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- OSWorld-Verified: 83% · Contributes. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Creative Writing v3: 1964.1 · Contributes. Source · Evaluator-reported · gpt-5.6-sol; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- MathArena ; ArXivMath Jun 2026: 86.73% · Contributes. Source · Archived source review · gpt-5.6-sol; max effort; MathArena June 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MathArena ; BrokenArXiv Jun 2026: 67.28% · Contributes. Source · Archived source review · gpt-5.6-sol; max effort; MathArena June 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- AIIQ Composite IQ: 136 · Not used. Source · Evaluator-reported · gpt-5.6-sol; max effort; AIIQ composite snapshot
Held-out composite, never a scoring input
- MCP Atlas †: 83.6% · Contributes. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- OfficeQA Pro †: 63.2% · Contributes. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- EQ-Bench 4: 1250.3 · Contributes. Source · Evaluator-reported · openai/gpt-5.6-sol; evaluator published configuration, effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CoWorkBench: 71.5% · Not used.
Not admitted after the September 6 model review: the collected 71.5 for GPT-5.6 Sol on CoWorkBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ²-Bench Telecom: 85.1% · Not used.
Not admitted after the September 6 model review: the collected 85.1 for GPT-5.6 Sol on τ²-Bench Telecom still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- PaperBench (Replication Score): 90.5% · Not used.
Not admitted after the September 6 model review: the collected 90.5 for GPT-5.6 Sol on PaperBench (Replication Score) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- QwenReactBench (Elo): 1564 · Not used.
Not admitted after the September 6 model review: the collected 1564 for GPT-5.6 Sol on QwenReactBench (Elo) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ERQA: 70% · Not used.
Not admitted after the September 6 model review: the collected 70 for GPT-5.6 Sol on ERQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vision2Web (Avg. Frontend/Webpage/etc.): 62.1% · Not used.
Not admitted after the September 6 model review: the collected 62.1 for GPT-5.6 Sol on Vision2Web (Avg. Frontend/Webpage/etc.) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MobileWorld: 76.9% · Not used.
Not admitted after the September 6 model review: the collected 76.9 for GPT-5.6 Sol on MobileWorld still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- WebArena-Verified: 69.7% · Not used.
Not admitted after the September 6 model review: the collected 69.7 for GPT-5.6 Sol on WebArena-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathArena ; ArXivLean Jun 2026: 37.5% · Contributes. Source · Archived source review · gpt-5.6-sol; max effort; MathArena June 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Frontier-Bench v0.1 ; Anthropic H2H: 34.4% · Contributes. Source · Archived source review · gpt-5.6-sol; max effort; Anthropic Opus 5 published head-to-head setup
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- FrontierCode v1.1 Main ; Anthropic H2H: 47.5% · Not used.
Not admitted after the September 6 model review: the collected 47.5 for GPT-5.6 Sol on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AutomationBench ; Anthropic H2H: 18.1% · Contributes. Source · Archived source review · gpt-5.6-sol; max effort; Anthropic Opus 5 published head-to-head setup
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Harvey Legal Agent Benchmark ; held-out: 2.5% · Not used.
Not admitted after the September 6 model review: the collected 2.5 for GPT-5.6 Sol on Harvey Legal Agent Benchmark ; held-out still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 44.329897% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 32.285714% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 4.0: 37.3% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol; max; Codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- GDP.pdf: 30.7% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (Max reasoning); Surge original evaluation
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 83% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 80.5% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 2108 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- OSWorld 2.0: 62.6% · Not used. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ARC-AGI-1: 97.5% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol; XHigh; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- ARC-AGI-3 (standard harness): 7.78% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- BenchCAD (Vision2Code subset): 83.3% · Not used. Source · Developer-reported · OpenAI full Vision2Code set; not Anthropic subset
System card identifies the OpenAI comparison as the full 17,900-file set, while this row is a modified 1000-file subset. Requires a separate protocol row.
- Terminal-Bench Science 0.1: 22.4% · Contributes. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- SRE-Bench (pass@1): 55.9% · Contributes. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- OSWorld 2.0 ; OpenAI H2H (offline subset): 65.7% · Contributes. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Agents' Last Exam ; OpenAI H2H: 53.6% · Contributes. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- ScreenSpot-Pro: 76.9% · Not used. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- HealthBench: 55.6% · Not used. Source · Developer-reported · GPT-5.6 Sol; OpenAI published HealthBench research evaluation; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- HealthBench Hard: 31.1% · Not used. Source · Developer-reported · GPT-5.6 Sol; OpenAI published HealthBench research evaluation; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- HealthBench Professional (length-adjusted): 60.5% · Contributes. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- GeneBench Pro: 32.3% · Not used. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- LifeSciBench: 59.9% · Not used. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- SEC-Bench Pro: 79.1% · Not used. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- ExploitGym: 30.3% · Not used. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Harvey LAB-AA: 87.18% · Not used.
Not admitted after the September 6 model review: the collected 87.18 for GPT-5.6 Sol on Harvey LAB-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EnterpriseOps-Gym-AA: 42.91% · Not used.
Not admitted after the September 6 model review: the collected 42.91 for GPT-5.6 Sol on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 26.111111% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-AnalystAgent: 47.5% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ITBench-AA: 56.214689% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Vending-Bench 2: 9619.368 · Contributes. Source · Evaluator-reported · GPT-5.6 Sol; Andon Labs Vending-Bench 2
Compare reasoning configurations using their mean over runs. Geometric mean, individual best runs, Fireworks Kimi endpoint, Vending Arena, and April DeepSeek Pro are not substituted.
- IFBench: 72.7% · Not used.
Not admitted after the September 6 model review: the collected 72.7 for GPT-5.6 Sol on IFBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AndroidWorld: 77.6% · Not used.
Not admitted after the September 6 model review: the collected 77.6 for GPT-5.6 Sol on AndroidWorld still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CWE-bench: 44.2% · Not used.
Not admitted after the September 6 model review: the collected 44.2 for GPT-5.6 Sol on CWE-bench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ZeroBench: 30% · Contributes. Source · Evaluator-reported · GPT-5.6 Sol (max); official main questions
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SimpleVQA: 66.6% · Not used.
Not admitted after the September 6 model review: the collected 66.6 for GPT-5.6 Sol on SimpleVQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathVision (with CI): 97.8% · Contributes. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; with Python tools
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- OmniDocBench 1.5: 85.8% · Contributes. Source · Developer-reported · GPT-5.6 Sol; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Internal Database Migration Tasks ; OpenAI: 42.7% · Contributes. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Internal Design Tasks ; OpenAI: 47.4% · Contributes. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Internal Data Science Tasks ; OpenAI: 30.5% · Contributes. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- OpenScore String Quartets (1 - OMR-NED): 0.19 · Not used. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- MedChemBench (internal): 47.4% · Not used. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- MRCR v2 (8 needles, 512K-1M): 73.8% · Not used. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- ExploitBench (Jun-Aug 2026, 300-turn limit): 5.5% · Not used. Source · Developer-reported · GPT-5.6 Sol; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- Mystery Game Puzzles v1.0.4 ; Epoch: 58% · Contributes. Source · Evaluator-reported · gpt-5.6-sol_max; task version 1.0.4
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- FrontierMath Erdos (68 conjectures) ; Epoch: 0 · Not used. Source · Evaluator-reported · gpt-5.6-sol_max; task version
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- Box Complex Work Eval (overall): 74% · Not used.
Not admitted after the September 6 model review: the collected 74 for GPT-5.6 Sol on Box Complex Work Eval (overall) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCyber (Sep 2026) ; Irregular: 34 · Not used.
Not admitted after the September 6 model review: the collected 34 for GPT-5.6 Sol on FrontierCyber (Sep 2026) ; Irregular still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CyScenarioBench (Sep 2026) ; Irregular: 27% · Not used.
Not admitted after the September 6 model review: the collected 27 for GPT-5.6 Sol on CyScenarioBench (Sep 2026) ; Irregular still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- BioMysteryBench v1 ; Vals: 71.11% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Tax Agent Bench v1 ; Vals: 67.95% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-sol; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- RuneBench (30m, mean ln(1 + XP/min)): 5.896324 · Contributes. Source · Archived source review · gpt56-xhigh
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 2284 · Contributes. Source · Archived source review · GPT-5.6 Sol (Max)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Roboflow Vision Evals (six-task mean, high tier): 81% · Contributes. Source · Archived source review · Require a complete six-task high-tier panel, the highest published benchmark tier. Retain high-tier scores even when lower than low-tier. Do not substitute the best of three repetitions: use the published mean. Each panel enters the fit once; do not also fit its components.
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1617 · Contributes. Source · Archived source review · gpt-5.6-sol-xhigh (codex-harness)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
Grok 4.6 · 27 used / 59 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
25 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±5.6 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1665.73 · Contributes. Source · Evaluator-reported · Grok 4.6 (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 88.389513% · Contributes. Source · Evaluator-reported · Grok 4.6 (high); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 67.477876% · Contributes. Source · Evaluator-reported · grok-4-6; medium; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1461.15072 · Contributes. Source · Evaluator-reported · grok-4.6-high; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 61 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 94.949495% · Contributes. Source · Evaluator-reported · Grok 4.6 (high); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 75% · Not used. Source · Evaluator-reported · Grok 4.6 (high)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 53.68% · Not used.
Not admitted after the September 6 model review: the collected 53.68 for Grok 4.6 High on Finance Agent v2 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Excel Modeling Benchmark ; Vals: 62.57% · Not used.
Not admitted after the September 6 model review: the collected 62.57 for Grok 4.6 High on Excel Modeling Benchmark ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SciCode: 56.481481% · Contributes. Source · Evaluator-reported · Grok 4.6 (high); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 44.068582% · Contributes. Source · Evaluator-reported · Grok 4.6 (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 48.08% · Not used.
Not admitted after the September 6 model review: the collected 48.08 for Grok 4.6 High on Legal Research Bench ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Harvey LAB (Vals): 15.8% · Not used.
Not admitted after the September 6 model review: the collected 15.8 for Grok 4.6 High on Harvey LAB (Vals) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Briefcase Overall Elo: 1549.11 · Contributes. Source · Evaluator-reported · Grok 4.6 (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 86.53% · Not used.
Not admitted after the September 6 model review: the collected 86.53 for Grok 4.6 High on MedScribe ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SWE-bench Verified ; Vals: 95.6% · Not used.
Not admitted after the September 6 model review: the collected 95.6 for Grok 4.6 High on SWE-bench Verified ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- LegalBench ; Vals: 86.31% · Not used.
Not admitted after the September 6 model review: the collected 86.31 for Grok 4.6 High on LegalBench ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- TaxEval v2 ; Vals: 71.1% · Not used.
Not admitted after the September 6 model review: the collected 71.1 for Grok 4.6 High on TaxEval v2 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Index: 71.1% · Not used.
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 66.85% · Not used.
Not admitted after the September 6 model review: the collected 66.85 for Grok 4.6 High on Public Benefits Bench ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MedCode ; Vals: 44.71% · Not used.
Not admitted after the September 6 model review: the collected 44.71 for Grok 4.6 High on MedCode ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 30.483333 · Contributes. Source · Evaluator-reported · Grok 4.6 (high); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 89.4% · Not used.
Not admitted after the September 6 model review: the collected 89.4 for Grok 4.6 High on MMLU Pro ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- LiveBench: 78.042% · Contributes. Source · Evaluator-reported · grok-4.6; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- MortgageTax ; Vals: 64.19% · Not used.
Not admitted after the September 6 model review: the collected 64.19 for Grok 4.6 High on MortgageTax ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SAGE ; Vals: 28.9% · Not used.
Not admitted after the September 6 model review: the collected 28.9 for Grok 4.6 High on SAGE ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Code Migration †: 44.57% · Not used.
Not admitted after the September 6 model review: the collected 44.57 for Grok 4.6 High on Code Migration † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Terminal-Bench v3.0: 26.5% · Contributes. Source · Evaluator-reported · Grok 4.6 (high); Grok Build
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveCodeBench: 88.22% · Not used.
Not admitted after the September 6 model review: the collected 88.22 for Grok 4.6 High on LiveCodeBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CyberGym: 79.7% · Not used.
Not admitted after the September 6 model review: the collected 79.7 for Grok 4.6 High on CyberGym still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SWE-Marathon v1.1: 31.9% · Not used.
Not admitted after the September 6 model review: the collected 31.9 for Grok 4.6 High on SWE-Marathon v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CursorBench v3.2: 70.8% · Contributes. Source · Evaluator-reported · Grok 4.6 Extra High; Cursor harness
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- APEX-Agents Mean Criteria Passed: 57.5% · Contributes. Source · Evaluator-reported · grok-4-6; high; loop_truncated_tools_agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- APEX-SWE (Pass@1, Terminus-2): 56.4% · Contributes. Source · Evaluator-reported · grok-4-6; high; terminus-2
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ARC-AGI-2: 67.1% · Not used.
Not admitted after the September 6 model review: the collected 67.1 for Grok 4.6 High on ARC-AGI-2 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SimpleQA: 53% · Not used.
Not admitted after the September 6 model review: the collected 53 for Grok 4.6 High on SimpleQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Extended: 61.31% · Contributes. Source · Evaluator-reported · Grok 4.6; high; grok-build
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- OfficeQA Pro †: 63.2% · Not used.
Not admitted after the September 6 model review: the collected 63.2 for Grok 4.6 High on OfficeQA Pro † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Main ; Anthropic H2H: 48% · Not used.
Not admitted after the September 6 model review: the collected 48 for Grok 4.6 High on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 50.721649% · Contributes. Source · Evaluator-reported · Grok 4.6 (high); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 19.714286% · Contributes. Source · Evaluator-reported · Grok 4.6 (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 4.0: 20.3% · Contributes. Source · Evaluator-reported · Grok 4.6; high; Grok Build
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- GDP.pdf: 17.2% · Contributes. Source · Evaluator-reported · Grok 4.6 (xHigh reasoning); Surge original evaluation
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 51% · Not used.
Not admitted after the September 6 model review: the collected 51 for Grok 4.6 High on ProofBench v1.1 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vibe Code Bench v1.1 ; Vals: 76.24% · Not used.
Not admitted after the September 6 model review: the collected 76.24 for Grok 4.6 High on Vibe Code Bench v1.1 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MineBench: 1931 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- ARC-AGI-1: 87.5% · Not used.
Not admitted after the September 6 model review: the collected 87.5 for Grok 4.6 High on ARC-AGI-1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ARC-AGI-3 (standard harness): 2.11% · Not used.
Not admitted after the September 6 model review: the collected 2.11 for Grok 4.6 High on ARC-AGI-3 (standard harness) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EnterpriseOps-Gym-AA: 48.34% · Not used.
Not admitted after the September 6 model review: the collected 48.34 for Grok 4.6 High on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 12.222222% · Contributes. Source · Evaluator-reported · Grok 4.6 (high); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-AnalystAgent: 41.25% · Contributes. Source · Evaluator-reported · Grok 4.6 (high); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Vending-Bench 2: 9047.0325 · Contributes. Source · Evaluator-reported · Grok 4.6; Andon Labs Vending-Bench 2
Compare reasoning configurations using their mean over runs. Geometric mean, individual best runs, Fireworks Kimi endpoint, Vending Arena, and April DeepSeek Pro are not substituted.
- CWE-bench: 38.2% · Contributes. Source · Evaluator-reported · Grok 4.6 High; high reasoning; grok-cli
Published high effort matches this model configuration. Other models pinned to max/xhigh are not approved from this high-only board.
- ZeroBench: 17% · Contributes. Source · Evaluator-reported · Grok 4.6 (xhigh); official main questions
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- BioMysteryBench v1 ; Vals: 72.222% · Not used.
Not admitted after the September 6 model review: the collected 72.222 for Grok 4.6 High on BioMysteryBench v1 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Tax Agent Bench v1 ; Vals: 70.788% · Not used.
Not admitted after the September 6 model review: the collected 70.788 for Grok 4.6 High on Tax Agent Bench v1 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- RuneBench (30m, mean ln(1 + XP/min)): 5.621351 · Contributes. Source · Archived source review · grok46-xhigh
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 2099 · Contributes. Source · Archived source review · Grok 4.6 (xHigh)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1625 · Contributes. Source · Archived source review · grok-4.6-high
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
Kimi K3 · 67 used / 93 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
62 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.0 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1587.5 · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 85.018727% · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 68.514412% · Contributes. Source · Evaluator-reported · kimi-k3; max; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1488.665886 · Contributes. Source · Evaluator-reported · kimi-k3-max; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 60 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 93.535354% · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 74.7% · Not used. Source · Evaluator-reported · Kimi K3 (max)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 54.4% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI Finance Agent v2 harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Excel Modeling Benchmark ; Vals: 66.4% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI Excel Modeling overall evaluation
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- SciCode: 59.490741% · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 46.895273% · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 44.23% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI Legal Research overall harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Harvey LAB (Vals): 10.83% · Not used.
Not admitted after the September 6 model review: the collected 10.83 for Kimi K3 on Harvey LAB (Vals) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Briefcase Overall Elo: 1499.88 · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 87.96% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI health evaluation
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- SWE-bench Verified ; Vals: 93.4% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI mini-SWE-agent bash harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- LegalBench ; Vals: 86.02% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI LegalBench evaluation
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- TaxEval v2 ; Vals: 75.72% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI max-compute evaluation
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- KingBench 3: 77.5% · Not used.
Not admitted after the September 6 model review: the collected 77.5 for Kimi K3 on KingBench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Index: 74.7% · Not used. Source · Evaluator-reported · kimi-k3; max effort; Vals AI index methodology
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 68.27% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI public-benefits evaluation
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MedCode ; Vals: 48.88% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI health evaluation
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- AutomationBench v1.0.6: 46.7% · Not used.
Not admitted after the September 6 model review: the collected 46.7 for Kimi K3 on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 19.7 · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 87.97% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI five-shot MMLU-Pro harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- LiveBench: 79.193048% · Contributes. Source · Evaluator-reported · kimi-k3; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- CorpFin v2 ; Vals: 71.56% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI CorpFin v2 evaluation with Sonnet 4.5 judge
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MortgageTax ; Vals: 66.34% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI max-compute evaluation
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- SAGE ; Vals: 54.26% · Contributes. Source · Archived source review · kimi-k3; max effort; Vals AI max-compute evaluation
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Code Migration †: 16.1% · Not used.
Not admitted after the September 6 model review: the collected 16.1 for Kimi K3 on Code Migration † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Toolathlon-Verified: 76.5% · Contributes. Source · Evaluator-reported · Kimi Kimi K3 (max) ✓; Default agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Agents' Last Exam: 28.3% · Contributes. Source · Evaluator-reported · Kimi K3; Max; Kimi Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Terminal-Bench v3.0: 17.4% · Not used.
Not admitted after the September 6 model review: the collected 17.4 for Kimi K3 on Terminal-Bench v3.0 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- LiveCodeBench: 87.19% · Not used.
Not admitted after the September 6 model review: the collected 87.19 for Kimi K3 on LiveCodeBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CyberGym: 80% · Contributes. Source · Evaluator-reported · Kimi K3; Z.ai GLM-5.3 model-card comparison table; GLM max, other comparator efforts not fully disclosed
CyberGym release-table percentage. Z.ai documents single-run pass@1 on 1,507 tasks, unlimited task timeout, Claude Code 2.1.207 and restricted network for GLM-5.3. Do not transfer those GLM-specific settings to every comparator. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- Humanity's Last Exam w/ tools: 59.8% · Contributes. Source · Evaluator-reported · Kimi K3; Z.ai GLM-5.3 model-card comparison table; GLM max, other comparator efforts not fully disclosed
HLE with tools release-table percentage. Z.ai documents temperature=1, top_p=.95, 163,840 output tokens, 300K context and GPT-5.6 Luna medium judge for its evaluation. Comparator-specific tools and judging settings are not fully disclosed. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- Design Arena (Elo): 1453 · Contributes. Source · Archived source review · kimi-k3; max effort; Design Arena 2026-08-11 rating snapshot
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Vals Multimodal Index: 73.42% · Not used. Source · Evaluator-reported · kimi-k3; max effort; Vals AI multimodal index methodology
Held-out composite, never a scoring input
- SWE-Marathon v1.1: 48.1% · Not used.
Not admitted after the September 6 model review: the collected 48.1 for Kimi K3 on SWE-Marathon v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 62.768% · Contributes. Source · Archived source review · kimi/kimi-k3 · Vals published row (effort not stated)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- FrontierSWE: 81.2% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierMath v2 Tier 4: 39.02439% · Contributes. Source · Evaluator-reported · kimi-k3_max; task version 2.0.0
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- CursorBench v3.2: 60.8% · Contributes. Source · Evaluator-reported · Kimi K3 Max; Cursor harness
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- APEX-Agents Mean Criteria Passed: 55.4% · Contributes. Source · Evaluator-reported · kimi-k3; max; loop_truncated_tools_agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CharXiv (with CI / RQ): 91.3% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; with Python tools
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- APEX-SWE (Pass@1, Terminus-2): 48% · Contributes. Source · Evaluator-reported · kimi-k3; max; terminus-2
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- BrowseComp: 91.2% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MMMU-Pro: 80.520231% · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ARC-AGI-2: 60.4% · Contributes. Source · Evaluator-reported · Kimi K3; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- OSWorld 2.0 ; Anthropic H2H: 58.3% · Contributes. Source · Archived source review · kimi-k3; max effort; Anthropic Opus 5 published head-to-head setup
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- ExploitBench: 32.2% · Not used.
Inherited comparator metric/scaffold equivalence to Astra Cap Percent is not established; the source flags possible contamination. Excluded pending protocol audit.
- PostTrainBench: 36.6% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierCode v1.1 Extended: 58.19% · Contributes. Source · Evaluator-reported · Kimi K3; none; mini-swe-agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- JobBench: 54.3% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- BabyVision (with CI): 85.7% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; with Python tools
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- PerceptionBench: 58.5% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- OSWorld-Verified: 84.8% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Creative Writing v3: 2070.8 · Contributes. Source · Evaluator-reported · kimi-k3; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- MathArena ; ArXivMath Jun 2026: 72.11% · Contributes. Source · Archived source review · kimi-k3; max effort; MathArena June 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MathArena ; BrokenArXiv Jun 2026: 51.85% · Contributes. Source · Archived source review · kimi-k3; max effort; MathArena June 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- AIIQ Composite IQ: 122 · Not used. Source · Evaluator-reported · kimi-k3; max effort; AIIQ composite snapshot
Held-out composite, never a scoring input
- MCP Atlas †: 84.2% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- OfficeQA Pro †: 63.3% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- EQ-Bench 4: 1339.3 · Contributes. Source · Evaluator-reported · moonshotai/kimi-k3; evaluator published configuration, effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- NL2Repo: 58% · Contributes. Source · Evaluator-reported · Kimi K3; Z.ai GLM-5.3 model-card comparison table; GLM max, other comparator efforts not fully disclosed
NL2Repo release-table percentage. Z.ai documents temperature=1, top_p=1, 64K output, 1M context and anti-cheating judges for its evaluation. Comparator-specific harness settings are not fully disclosed. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- τ²-Bench Telecom: 80.63% · Not used.
Not admitted after the September 6 model review: the collected 80.63 for Kimi K3 on τ²-Bench Telecom still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Main ; Anthropic H2H: 44.2% · Not used.
Not admitted after the September 6 model review: the collected 44.2 for Kimi K3 on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathArena ; ArXivMath May 2026: 61.67% · Contributes. Source · Archived source review · kimi-k3; max effort; MathArena May 2026 competition harness
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- τ³-Banking: 45.979381% · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 23.428571% · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- GDP.pdf: 19% · Contributes. Source · Evaluator-reported · Kimi K3 (Max reasoning); Surge original evaluation
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 87% · Not used.
Not admitted after the September 6 model review: the collected 87 for Kimi K3 on ProofBench v1.1 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vibe Code Bench v1.1 ; Vals: 84.96% · Not used.
Not admitted after the September 6 model review: the collected 84.96 for Kimi K3 on Vibe Code Bench v1.1 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MineBench: 1812 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- ARC-AGI-1: 94.5% · Contributes. Source · Evaluator-reported · Kimi K3; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- Harvey LAB-AA: 94.64% · Not used.
Not admitted after the September 6 model review: the collected 94.64 for Kimi K3 on Harvey LAB-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EnterpriseOps-Gym-AA: 45.33% · Not used.
Not admitted after the September 6 model review: the collected 45.33 for Kimi K3 on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 38.333333% · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-AnalystAgent: 38.75% · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ITBench-AA: 47.693032% · Contributes. Source · Evaluator-reported · Kimi K3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSearchQA: 95% · Not used. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vending-Bench 2: 5165.035 · Contributes. Source · Evaluator-reported · Kimi K3 (Moonshot); Andon Labs Vending-Bench 2
Compare reasoning configurations using their mean over runs. Geometric mean, individual best runs, Fireworks Kimi endpoint, Vending Arena, and April DeepSeek Pro are not substituted.
- IFBench: 74.4% · Not used.
Not admitted after the September 6 model review: the collected 74.4 for Kimi K3 on IFBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CWE-bench: 22.5% · Not used.
Not admitted after the September 6 model review: the collected 22.5 for Kimi K3 on CWE-bench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ZeroBench: 23% · Not used.
Not admitted after the September 6 model review: the collected 23 for Kimi K3 on ZeroBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathVision (with CI): 97.8% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; with Python tools
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- OmniDocBench 1.5: 91.1% · Contributes. Source · Developer-reported · Kimi K3; max; Kimi developer comparison table; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Mystery Game Puzzles v1.0.4 ; Epoch: 26% · Contributes. Source · Evaluator-reported · kimi-k3_max; task version 1.0.4
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- BioMysteryBench v1 ; Vals: 71.481% · Not used.
Not admitted after the September 6 model review: the collected 71.481 for Kimi K3 on BioMysteryBench v1 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Tax Agent Bench v1 ; Vals: 68.672% · Not used.
Not admitted after the September 6 model review: the collected 68.672 for Kimi K3 on Tax Agent Bench v1 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- RuneBench (30m, mean ln(1 + XP/min)): 3.753577 · Contributes. Source · Archived source review · kimi3
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 1932 · Contributes. Source · Archived source review · Kimi K3 (Max)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1674 · Contributes. Source · Archived source review · kimi-k3-max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
GLM-5.3 · 42 used / 56 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
39 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.3 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1679.47 · Contributes. Source · Evaluator-reported · GLM-5.3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 83.895131% · Contributes. Source · Evaluator-reported · GLM-5.3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 68.957871% · Contributes. Source · Evaluator-reported · glm-5-3; max; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1482.03722 · Contributes. Source · Evaluator-reported · glm-5.3-max; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 60 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 91.717172% · Contributes. Source · Evaluator-reported · GLM-5.3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 76.33% · Not used. Source · Evaluator-reported · GLM-5.3 (max)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 55.84% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 56.34% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 59.027778% · Contributes. Source · Evaluator-reported · GLM-5.3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 42.261353% · Contributes. Source · Evaluator-reported · GLM-5.3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 49.04% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 8.33% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1518.58 · Contributes. Source · Evaluator-reported · GLM-5.3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 88.81% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 95.4% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 84.84% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 72.36% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- KingBench 3: 91.25% · Not used.
Not admitted after the September 6 model review: the collected 91.25 for GLM-5.3 on KingBench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Public Benefits Bench ; Vals: 68.54% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 42.86% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AutomationBench v1.0.6: 48.2% · Not used.
Not admitted after the September 6 model review: the collected 48.2 for GLM-5.3 on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 14.3 · Contributes. Source · Evaluator-reported · GLM-5.3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 86.77% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 76.13781% · Contributes. Source · Evaluator-reported · glm-5.3; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- Code Migration †: 44.22% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Toolathlon-Verified: 73% · Not used.
Not admitted after the September 6 model review: the collected 73 for GLM-5.3 on Toolathlon-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Agents' Last Exam: 28.5% · Not used.
Not admitted after the September 6 model review: the collected 28.5 for GLM-5.3 on Agents' Last Exam still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Terminal-Bench v3.0: 32.4% · Contributes. Source · Evaluator-reported · GLM-5.3 (max); Claude Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveCodeBench: 80.53% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CyberGym: 84.5% · Contributes. Source · Evaluator-reported · GLM-5.3; Z.ai GLM-5.3 model-card comparison table; GLM max, other comparator efforts not fully disclosed
CyberGym release-table percentage. Z.ai documents single-run pass@1 on 1,507 tasks, unlimited task timeout, Claude Code 2.1.207 and restricted network for GLM-5.3. Do not transfer those GLM-specific settings to every comparator. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- Humanity's Last Exam w/ tools: 62.5% · Contributes. Source · Evaluator-reported · GLM-5.3; Z.ai GLM-5.3 model-card comparison table; GLM max, other comparator efforts not fully disclosed
HLE with tools release-table percentage. Z.ai documents temperature=1, top_p=.95, 163,840 output tokens, 300K context and GPT-5.6 Luna medium judge for its evaluation. Comparator-specific tools and judging settings are not fully disclosed. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- SWE-Marathon v1.1: 42.5% · Not used.
Not admitted after the September 6 model review: the collected 42.5 for GLM-5.3 on SWE-Marathon v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 66.252% · Contributes. Source · Archived source review · zai/glm-5.3 · reasoning_effort=max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- SkillsBench †: 47.51% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierSWE: 78.1% · Not used.
Not admitted after the September 6 model review: the collected 78.1 for GLM-5.3 on FrontierSWE still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierMath v2 Tier 4: 29.27% · Not used.
Not admitted after the September 6 model review: the collected 29.27 for GLM-5.3 on FrontierMath v2 Tier 4 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ExploitBench: 54.4% · Not used.
Inherited comparator metric/scaffold equivalence to Astra Cap Percent is not established; the source flags possible contamination. Excluded pending protocol audit.
- PostTrainBench: 39.8% · Not used.
Not admitted after the September 6 model review: the collected 39.8 for GLM-5.3 on PostTrainBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Creative Writing v3: 2062.4 · Contributes. Source · Evaluator-reported · *GLM-5.3; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- NL2Repo: 58% · Contributes. Source · Evaluator-reported · GLM-5.3; Z.ai GLM-5.3 model-card comparison table; GLM max, other comparator efforts not fully disclosed
NL2Repo release-table percentage. Z.ai documents temperature=1, top_p=1, 64K output, 1M context and anti-cheating judges for its evaluation. Comparator-specific harness settings are not fully disclosed. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- τ³-Banking: 50.309278% · Contributes. Source · Evaluator-reported · GLM-5.3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 19.142857% · Contributes. Source · Evaluator-reported · GLM-5.3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 4.0: 41.8% · Contributes. Source · Evaluator-reported · GLM-5.3; max; Claude Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 49% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 78.13% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 1850 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- EnterpriseOps-Gym-AA: 36.44% · Not used.
Not admitted after the September 6 model review: the collected 36.44 for GLM-5.3 on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 48.333333% · Contributes. Source · Evaluator-reported · GLM-5.3 (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Vending-Bench 2: 8163.61 · Contributes. Source · Evaluator-reported · GLM-5.3; Andon Labs Vending-Bench 2
Compare reasoning configurations using their mean over runs. Geometric mean, individual best runs, Fireworks Kimi endpoint, Vending Arena, and April DeepSeek Pro are not substituted.
- CWE-bench: 31.1% · Contributes. Source · Evaluator-reported · GLM-5.3; high; opencode
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Mystery Game Puzzles v1.0.4 ; Epoch: 33% · Not used.
Not admitted after the September 6 model review: the collected 33 for GLM-5.3 on Mystery Game Puzzles v1.0.4 ; Epoch still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Tax Agent Bench v1 ; Vals: 73.09% · Contributes. Source · Evaluator-reported · zai_glm-5.3; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- RuneBench (30m, mean ln(1 + XP/min)): 5.001589 · Contributes. Source · Archived source review · glm53
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 1786 · Contributes. Source · Archived source review · GLM-5.3 (Max)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1609 · Contributes. Source · Archived source review · glm-5.3-max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
Gemini 3.8 Flash · 40 used / 52 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
39 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.3 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1464.96 · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 87.640449% · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 73.825503% · Contributes. Source · Evaluator-reported · gemini-3-8-flash; high; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1494 · Not used.
Not admitted after the September 6 model review: the collected 1494 for Gemini 3.8 Flash on LMArena (Chatbot Arena) ; Text still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA Intelligence Index v4.1.1 (score): 59 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 95.252525% · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 82% · Not used. Source · Evaluator-reported · Gemini 3.8 Flash (high)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 61.44% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 72.2% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 56.597222% · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 47.822057% · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 38.94% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 10% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1201.13 · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash (high); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 84.5% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 80% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 86.99% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 74.45% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vals Index: 62.25% · Not used.
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 65.29% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 48.13% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Omniscience: 29.55 · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 90.22% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 75.83% · Not used.
Not admitted after the September 6 model review: the collected 75.83 for Gemini 3.8 Flash on LiveBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MortgageTax ; Vals: 65.34% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 35.06% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 36.55% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveCodeBench: 89.48% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 71.899% · Contributes. Source · Archived source review · google/gemini-3.8-flash · reasoning_effort=high
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- SkillsBench †: 57.98% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CursorBench v3.2: 69.2% · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash High; Cursor harness
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Pro: 61.6% · Not used.
Not admitted after the September 6 model review: the collected 61.6 for Gemini 3.8 Flash on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CharXiv (with CI / RQ): 86.2% · Not used.
Not admitted after the September 6 model review: the collected 86.2 for Gemini 3.8 Flash on CharXiv (with CI / RQ) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MMMU-Pro: 85.606936% · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- FrontierCode v1.1 Extended: 53.45% · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash; medium; chisel
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LVBench (with Memory): 87.8% · Not used.
Not admitted after the September 6 model review: the collected 87.8 for Gemini 3.8 Flash on LVBench (with Memory) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathArena ; ArXivMath Jun 2026: 68.03% · Not used.
Not admitted after the September 6 model review: the collected 68.03 for Gemini 3.8 Flash on MathArena ; ArXivMath Jun 2026 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathArena ; BrokenArXiv Jun 2026: 20.83% · Not used.
Not admitted after the September 6 model review: the collected 20.83 for Gemini 3.8 Flash on MathArena ; BrokenArXiv Jun 2026 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Main ; Anthropic H2H: 43.6% · Not used.
Not admitted after the September 6 model review: the collected 43.6 for Gemini 3.8 Flash on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 45.773196% · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash (medium); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 18.285714% · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 4.0: 19.1% · Contributes. Source · Evaluator-reported · Gemini 3.8 Flash; high; mini-SWE-agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 48% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 78.65% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 1917 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- HealthBench Professional (length-adjusted): 52.1% · Contributes. Source · Developer-reported · Gemini 3.8 Flash; publisher-reported maximum across reasoning efforts; OpenAI research/API evaluation
Developer-reported maximum across reasoning modes. Exact selected mode is not individually disclosed. Research prompts/tools may differ from production; protocol-specific rows remain separate.
- BioMysteryBench v1 ; Vals: 62.22% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Tax Agent Bench v1 ; Vals: 66.77% · Contributes. Source · Evaluator-reported · google_gemini-3.8-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- RuneBench (30m, mean ln(1 + XP/min)): 5.962732 · Contributes. Source · Archived source review · gemini38flash
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 1811 · Contributes. Source · Archived source review · Gemini 3.8 Flash
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Roboflow Vision Evals (six-task mean, high tier): 86.516667% · Contributes. Source · Archived source review · Require a complete six-task high-tier panel, the highest published benchmark tier. Retain high-tier scores even when lower than low-tier. Do not substitute the best of three repetitions: use the published mean. Each panel enters the fit once; do not also fit its components.
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1567 · Contributes. Source · Archived source review · gemini-3.8-flash-high
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
Gemini 3.7 Flash · 41 used / 62 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
39 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.3 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1435.34 · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 85.76779% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 65.486726% · Contributes. Source · Evaluator-reported · gemini-3-7-flash; medium; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1490 · Not used.
Not admitted after the September 6 model review: the collected 1490 for Gemini 3.7 Flash on LMArena (Chatbot Arena) ; Text still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA Intelligence Index v4.1.1 (score): 56 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 94.545455% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 80% · Not used. Source · Evaluator-reported · Gemini 3.7 Flash (high)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 59.04% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 71.33% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 59.837963% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (medium); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 47.868397% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 34.62% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 8.75% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1116.43 · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (high); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 83.94% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 80.8% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 87.26% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 74.73% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vals Index: 59.31% · Not used. Source · Evaluator-reported · gemini-3.7-flash; high effort; Vals AI index methodology
Held-out composite, never a scoring input
- MedCode ; Vals: 53.39% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AutomationBench v1.0.6: 30.44% · Not used.
Not admitted after the September 6 model review: the collected 30.44 for Gemini 3.7 Flash on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 26.483333 · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 90.12% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 78.8% · Not used.
Not admitted after the September 6 model review: the collected 78.8 for Gemini 3.7 Flash on LiveBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MortgageTax ; Vals: 66.65% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 49.23% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 34.8% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Agents' Last Exam: 26.3% · Not used.
Not admitted after the September 6 model review: the collected 26.3 for Gemini 3.7 Flash on Agents' Last Exam still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Terminal-Bench v3.0: 14.9% · Not used.
Not admitted after the September 6 model review: the collected 14.9 for Gemini 3.7 Flash on Terminal-Bench v3.0 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- LiveCodeBench: 88.65% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-Marathon v1.1: 22.5% · Not used.
Not admitted after the September 6 model review: the collected 22.5 for Gemini 3.7 Flash on SWE-Marathon v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 68.663% · Contributes. Source · Archived source review · google/gemini-3.7-flash · reasoning_effort=high
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- FrontierMath v2 Tier 4: 36.59% · Not used.
Not admitted after the September 6 model review: the collected 36.59 for Gemini 3.7 Flash on FrontierMath v2 Tier 4 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CursorBench v3.2: 61.6% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash High; Cursor harness
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Pro: 60.4% · Not used.
Not admitted after the September 6 model review: the collected 60.4 for Gemini 3.7 Flash on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CharXiv (with CI / RQ): 88.7% · Not used.
Not admitted after the September 6 model review: the collected 88.7 for Gemini 3.7 Flash on CharXiv (with CI / RQ) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MMMU-Pro: 85.491329% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ARC-AGI-2: 84.6% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash; High; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- LVBench (with Memory): 85.4% · Not used.
Not admitted after the September 6 model review: the collected 85.4 for Gemini 3.7 Flash on LVBench (with Memory) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Creative Writing v3: 1725.7 · Not used.
Not admitted after the September 6 model review: the collected 1725.7 for Gemini 3.7 Flash on Creative Writing v3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathArena ; ArXivMath Jun 2026: 63.27% · Not used.
Not admitted after the September 6 model review: the collected 63.27 for Gemini 3.7 Flash on MathArena ; ArXivMath Jun 2026 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathArena ; BrokenArXiv Jun 2026: 11.11% · Not used.
Not admitted after the September 6 model review: the collected 11.11 for Gemini 3.7 Flash on MathArena ; BrokenArXiv Jun 2026 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Main ; Anthropic H2H: 43.6% · Not used.
Not admitted after the September 6 model review: the collected 43.6 for Gemini 3.7 Flash on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 35.463918% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (medium); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 14.285714% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (high); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 4.0: 11.2% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash; high; mini-SWE-agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- GDP.pdf: 23.8% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (High reasoning); Surge original evaluation
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 58% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 70.39% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 1867 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- ARC-AGI-1: 95.5% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash; High; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- Harvey LAB-AA: 90.66% · Not used.
Not admitted after the September 6 model review: the collected 90.66 for Gemini 3.7 Flash on Harvey LAB-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EnterpriseOps-Gym-AA: 50.4% · Not used.
Not admitted after the September 6 model review: the collected 50.4 for Gemini 3.7 Flash on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 15% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (high); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-AnalystAgent: 60% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash (high); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CWE-bench: 44% · Contributes. Source · Evaluator-reported · Gemini 3.7 Flash; high reasoning; Antigravity
Published high effort matches this model configuration. Other models pinned to max/xhigh are not approved from this high-only board.
- Mystery Game Puzzles v1.0.4 ; Epoch: 37% · Not used.
Not admitted after the September 6 model review: the collected 37 for Gemini 3.7 Flash on Mystery Game Puzzles v1.0.4 ; Epoch still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Tax Agent Bench v1 ; Vals: 57.66% · Contributes. Source · Evaluator-reported · google_gemini-3.7-flash; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- RuneBench (30m, mean ln(1 + XP/min)): 5.963375 · Contributes. Source · Archived source review · gemini37flash
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 1854 · Contributes. Source · Archived source review · Gemini 3.7 Flash
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Roboflow Vision Evals (six-task mean, high tier): 86.216667% · Contributes. Source · Archived source review · Require a complete six-task high-tier panel, the highest published benchmark tier. Retain high-tier scores even when lower than low-tier. Do not substitute the best of three repetitions: use the published mean. Each panel enters the fit once; do not also fit its components.
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1587 · Contributes. Source · Archived source review · gemini-3.7-flash-high
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
GPT-5.6 Terra · 54 used / 80 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
50 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.2 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1482.49 · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (xhigh); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 88.014981% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 69.62306% · Contributes. Source · Evaluator-reported · gpt-5-6-terra; max; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1466.420286 · Contributes. Source · Evaluator-reported · gpt-5.6-terra-xhigh; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 57 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 92.525253% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 79.7% · Not used. Source · Evaluator-reported · GPT-5.6 Terra (max)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 54.44% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 66.2% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 54.976852% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); Artificial Analysis published evaluation harness
Best published result among 3 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 42.910102% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 41.35% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 0.83% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1341.29 · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (xhigh); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 82.87% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 95.4% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 85.11% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 76.17% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- KingBench 3: 62.9% · Not used.
Not admitted after the September 6 model review: the collected 62.9 for GPT-5.6 Terra on KingBench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Index: 56.53% · Not used.
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 62.38% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 43.41% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AutomationBench v1.0.6: 37.17% · Not used.
Not admitted after the September 6 model review: the collected 37.17 for GPT-5.6 Terra on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 0.05 · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 86.66% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 77.93625% · Contributes. Source · Evaluator-reported · gpt-5.6-terra-max; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- CorpFin v2 ; Vals: 65.31% · Not used.
Not admitted after the September 6 model review: the collected 65.31 for GPT-5.6 Terra on CorpFin v2 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MortgageTax ; Vals: 67.33% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 47% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 47.8% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Toolathlon-Verified: 74.9% · Not used.
Not admitted after the September 6 model review: the collected 74.9 for GPT-5.6 Terra on Toolathlon-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Agents' Last Exam: 28% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra; Max; Codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Terminal-Bench v3.0: 20.8% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); Codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveCodeBench: 85.93% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CyberGym: 81.8% · Not used.
Not admitted after the September 6 model review: the collected 81.8 for GPT-5.6 Terra on CyberGym still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Design Arena (Elo): 1280 · Not used.
Not admitted after the September 6 model review: the collected 1280 for GPT-5.6 Terra on Design Arena (Elo) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Multimodal Index: 65.07% · Not used.
Held-out composite, never a scoring input
- SWE-Marathon v1.1: 32.5% · Not used.
Not admitted after the September 6 model review: the collected 32.5 for GPT-5.6 Terra on SWE-Marathon v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 72.345% · Contributes. Source · Archived source review · openai/gpt-5.6-terra · reasoning_effort=max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- SkillsBench †: 58.9% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierMath v2 Tier 4: 70.731707% · Contributes. Source · Evaluator-reported · gpt-5.6-terra_max; task version 2.0.0
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- CursorBench v3.2: 64.9% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra Max; Cursor harness
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Pro: 63.4% · Not used.
Not admitted after the September 6 model review: the collected 63.4 for GPT-5.6 Terra on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CharXiv (with CI / RQ): 88% · Not used.
Not admitted after the September 6 model review: the collected 88 for GPT-5.6 Terra on CharXiv (with CI / RQ) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- BrowseComp: 87.5% · Not used.
Not admitted after the September 6 model review: the collected 87.5 for GPT-5.6 Terra on BrowseComp still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MMMU-Pro: 80.693642% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ARC-AGI-2: 83.9% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- OSWorld 2.0 ; Anthropic H2H: 50.2% · Not used.
Not admitted after the September 6 model review: the collected 50.2 for GPT-5.6 Terra on OSWorld 2.0 ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ExploitBench: 52.9% · Not used.
Inherited comparator metric/scaffold equivalence to Astra Cap Percent is not established; the source flags possible contamination. Excluded pending protocol audit.
- SimpleQA: 43.1% · Not used.
Not admitted after the September 6 model review: the collected 43.1 for GPT-5.6 Terra on SimpleQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Extended: 55.84% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra; max; codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Creative Writing v3: 1850 · Contributes. Source · Evaluator-reported · gpt-5.6-terra; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- AIIQ Composite IQ: 132 · Not used.
Held-out composite, never a scoring input
- EQ-Bench 4: 1234 · Contributes. Source · Evaluator-reported · openai/gpt-5.6-terra; evaluator published configuration, effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- τ²-Bench Telecom: 86.3% · Not used.
Not admitted after the September 6 model review: the collected 86.3 for GPT-5.6 Terra on τ²-Bench Telecom still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Main ; Anthropic H2H: 41.3% · Not used.
Not admitted after the September 6 model review: the collected 41.3 for GPT-5.6 Terra on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 40.206186% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 30% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 4.0: 21.5% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra; max; Codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- GDP.pdf: 24.7% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (Medium reasoning); Surge original evaluation
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 74% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 74.59% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ARC-AGI-1: 96.5% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- ARC-AGI-3 (standard harness): 0.8% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- HealthBench: 58.7% · Not used. Source · Developer-reported · GPT-5.6 Terra; OpenAI published HealthBench research evaluation; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- HealthBench Hard: 34.3% · Not used. Source · Developer-reported · GPT-5.6 Terra; OpenAI published HealthBench research evaluation; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- HealthBench Professional (length-adjusted): 57.7% · Not used.
Source selects the maximum score across efforts; exact highest-effort result is not established.
- Harvey LAB-AA: 85.18% · Not used.
Not admitted after the September 6 model review: the collected 85.18 for GPT-5.6 Terra on Harvey LAB-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EnterpriseOps-Gym-AA: 38.5% · Not used.
Not admitted after the September 6 model review: the collected 38.5 for GPT-5.6 Terra on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 31.666667% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ITBench-AA: 51.035782% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Vending-Bench 2: 7343.206 · Contributes. Source · Evaluator-reported · GPT-5.6 Terra; Andon Labs Vending-Bench 2
Compare reasoning configurations using their mean over runs. Geometric mean, individual best runs, Fireworks Kimi endpoint, Vending Arena, and April DeepSeek Pro are not substituted.
- IFBench: 71.2% · Not used.
Not admitted after the September 6 model review: the collected 71.2 for GPT-5.6 Terra on IFBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ZeroBench: 19% · Contributes. Source · Evaluator-reported · GPT-5.6 Terra (max); official main questions
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Mystery Game Puzzles v1.0.4 ; Epoch: 35% · Contributes. Source · Evaluator-reported · gpt-5.6-terra_max; task version 1.0.4
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- Tax Agent Bench v1 ; Vals: 65.2% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-terra; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- RuneBench (30m, mean ln(1 + XP/min)): 5.881388 · Contributes. Source · Archived source review · gpt56terra-xhigh
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 1963 · Contributes. Source · Archived source review · GPT-5.6-Terra
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Roboflow Vision Evals (six-task mean, high tier): 74.25% · Contributes. Source · Archived source review · Require a complete six-task high-tier panel, the highest published benchmark tier. Retain high-tier scores even when lower than low-tier. Do not substitute the best of three repetitions: use the published mean. Each panel enters the fit once; do not also fit its components.
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1520 · Contributes. Source · Archived source review · gpt-5.6-terra-xhigh (codex-harness)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
Qwen3.8-Max · 46 used / 78 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
45 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.2 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1633.53 · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 81.273408% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 57.461024% · Contributes. Source · Evaluator-reported · qwen3-8-max; xhigh; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1479.556758 · Contributes. Source · Evaluator-reported · qwen3.8-max; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 58 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 92.727273% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 74.3% · Not used. Source · Evaluator-reported · Qwen3.8 Max
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 50.59% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 60.07% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 53.240741% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 43.04912% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 47.6% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 10.42% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1392.33 · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 84.95% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 85.6% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 83.61% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 75.55% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- KingBench 3: 81.25% · Not used.
Not admitted after the September 6 model review: the collected 81.25 for Qwen3.8 Max on KingBench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Index: 66.12% · Not used. Source · Evaluator-reported · qwen3.8-max; max effort; Vals AI index methodology
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 67.12% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 40.67% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AutomationBench v1.0.6: 39.8% · Not used.
Not admitted after the September 6 model review: the collected 39.8 for Qwen3.8 Max on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 3.4 · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 88.6% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 78.5% · Not used.
Not admitted after the September 6 model review: the collected 78.5 for Qwen3.8 Max on LiveBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CorpFin v2 ; Vals: 65.85% · Contributes. Source · Archived source review · qwen3.8-max; max effort; Vals AI CorpFin v2 evaluation with Sonnet 4.5 judge
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MortgageTax ; Vals: 63.99% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 51.25% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 23.96% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Toolathlon-Verified: 72.5% · Not used.
Not admitted after the September 6 model review: the collected 72.5 for Qwen3.8 Max on Toolathlon-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Agents' Last Exam: 27% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; XHigh; Claude Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveCodeBench: 87.85% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CyberGym: 78.5% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Z.ai GLM-5.3 model-card comparison table; GLM max, other comparator efforts not fully disclosed
CyberGym release-table percentage. Z.ai documents single-run pass@1 on 1,507 tasks, unlimited task timeout, Claude Code 2.1.207 and restricted network for GLM-5.3. Do not transfer those GLM-specific settings to every comparator. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- Humanity's Last Exam w/ tools: 56.2% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Z.ai GLM-5.3 model-card comparison table; GLM max, other comparator efforts not fully disclosed
HLE with tools release-table percentage. Z.ai documents temperature=1, top_p=.95, 163,840 output tokens, 300K context and GPT-5.6 Luna medium judge for its evaluation. Comparator-specific tools and judging settings are not fully disclosed. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- Design Arena (Elo): 1388 · Contributes. Source · Archived source review · qwen3.8-max; max effort; Design Arena 2026-08-11 rating snapshot
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Vals Multimodal Index: 65.39% · Not used. Source · Evaluator-reported · qwen3.8-max; max effort; Vals AI multimodal index methodology
Held-out composite, never a scoring input
- ProgramBench v1 (Raw Pass Rate) ; Vals: 39.971% · Contributes. Source · Archived source review · alibaba/qwen3.8-max · Vals published row (effort not stated)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- SkillsBench †: 42.01% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierSWE: 73.5% · Not used.
Not admitted after the September 6 model review: the collected 73.5 for Qwen3.8 Max on FrontierSWE still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierMath v2 Tier 4: 46.3% · Not used.
Not admitted after the September 6 model review: the collected 46.3 for Qwen3.8 Max on FrontierMath v2 Tier 4 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SWE-bench Pro: 67.7% · Not used.
Not admitted after the September 6 model review: the collected 67.7 for Qwen3.8 Max on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CharXiv (with CI / RQ): 93.5% · Not used.
Not admitted after the September 6 model review: the collected 93.5 for Qwen3.8 Max on CharXiv (with CI / RQ) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MMMU-Pro: 82.312139% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- OSWorld 2.0 ; Anthropic H2H: 46.7% · Contributes. Source · Archived source review · qwen3.8-max; max effort; Anthropic Opus 5 published head-to-head setup
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- ExploitBench: 28.8% · Not used.
Inherited comparator metric/scaffold equivalence to Astra Cap Percent is not established; the source flags possible contamination. Excluded pending protocol audit.
- JobBench: 53.4% · Not used.
Not admitted after the September 6 model review: the collected 53.4 for Qwen3.8 Max on JobBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- BabyVision (with CI): 91.3% · Not used.
Not admitted after the September 6 model review: the collected 91.3 for Qwen3.8 Max on BabyVision (with CI) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- PerceptionBench: 63.5% · Not used.
Not admitted after the September 6 model review: the collected 63.5 for Qwen3.8 Max on PerceptionBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- LVBench (with Memory): 85.6% · Not used.
Not admitted after the September 6 model review: the collected 85.6 for Qwen3.8 Max on LVBench (with Memory) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- OSWorld-Verified: 86.1% · Not used.
Not admitted after the September 6 model review: the collected 86.1 for Qwen3.8 Max on OSWorld-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Creative Writing v3: 1842 · Contributes. Source · Evaluator-reported · *Qwen/Qwen3.8-2.4T-A95B; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- NL2Repo: 55.9% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Z.ai GLM-5.3 model-card comparison table; GLM max, other comparator efforts not fully disclosed
NL2Repo release-table percentage. Z.ai documents temperature=1, top_p=1, 64K output, 1M context and anti-cheating judges for its evaluation. Comparator-specific harness settings are not fully disclosed. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- CoWorkBench: 74.8% · Not used.
Not admitted after the September 6 model review: the collected 74.8 for Qwen3.8 Max on CoWorkBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- PaperBench (Replication Score): 93% · Not used.
Not admitted after the September 6 model review: the collected 93 for Qwen3.8 Max on PaperBench (Replication Score) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- QwenReactBench (Elo): 1724 · Not used.
Not admitted after the September 6 model review: the collected 1724 for Qwen3.8 Max on QwenReactBench (Elo) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ERQA: 77.8% · Not used.
Not admitted after the September 6 model review: the collected 77.8 for Qwen3.8 Max on ERQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vision2Web (Avg. Frontend/Webpage/etc.): 69% · Not used.
Not admitted after the September 6 model review: the collected 69 for Qwen3.8 Max on Vision2Web (Avg. Frontend/Webpage/etc.) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MobileWorld: 77.8% · Not used.
Not admitted after the September 6 model review: the collected 77.8 for Qwen3.8 Max on MobileWorld still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- WebArena-Verified: 66.8% · Not used.
Not admitted after the September 6 model review: the collected 66.8 for Qwen3.8 Max on WebArena-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 51.340206% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 20% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- GDP.pdf: 23.2% · Contributes. Source · Evaluator-reported · Qwen 3.8 Max (xHigh reasoning); Surge original evaluation
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 58% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 64.7% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-max; Vals sole published default; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 1333 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- MLCR-AA (Medical Long Context Reasoning): 19.444444% · Contributes. Source · Evaluator-reported · Qwen3.8 Max; Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- IFBench: 82.8% · Not used.
Not admitted after the September 6 model review: the collected 82.8 for Qwen3.8 Max on IFBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AndroidWorld: 85.3% · Not used.
Not admitted after the September 6 model review: the collected 85.3 for Qwen3.8 Max on AndroidWorld still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CWE-bench: 37.5% · Contributes. Source · Evaluator-reported · Qwen3.8-Max; high; opencode
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ZeroBench: 24% · Not used.
Not admitted after the September 6 model review: the collected 24 for Qwen3.8 Max on ZeroBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SimpleVQA: 75% · Not used.
Not admitted after the September 6 model review: the collected 75 for Qwen3.8 Max on SimpleVQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathVision (with CI): 97.7% · Not used.
Not admitted after the September 6 model review: the collected 97.7 for Qwen3.8 Max on MathVision (with CI) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- OmniDocBench 1.5: 92.1% · Not used.
Not admitted after the September 6 model review: the collected 92.1 for Qwen3.8 Max on OmniDocBench 1.5 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- RuneBench (30m, mean ln(1 + XP/min)): 5.347003 · Contributes. Source · Archived source review · qwen38max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 1847 · Contributes. Source · Archived source review · Qwen3.8-Max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Roboflow Vision Evals (six-task mean, high tier): 85% · Contributes. Source · Archived source review · Require a complete six-task high-tier panel, the highest published benchmark tier. Retain high-tier scores even when lower than low-tier. Do not substitute the best of three repetitions: use the published mean. Each panel enters the fit once; do not also fit its components.
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1670 · Contributes. Source · Archived source review · qwen3.8-max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
Muse Spark 1.2 · 36 used / 52 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
36 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.4 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1527.08 · Contributes. Source · Evaluator-reported · Muse Spark 1.2 (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 80.149813% · Contributes. Source · Evaluator-reported · Muse Spark 1.2 (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 54.867257% · Contributes. Source · Evaluator-reported · muse-spark-1-2; xhigh; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1499 · Not used.
Not admitted after the September 6 model review: the collected 1499 for Muse Spark 1.2 on LMArena (Chatbot Arena) ; Text still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA Intelligence Index v4.1.1 (score): 57 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 90.40404% · Contributes. Source · Evaluator-reported · Muse Spark 1.2 (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 83.3% · Not used. Source · Evaluator-reported · Muse Spark 1.2 (xhigh)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 60.6% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 56.98% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 57.407407% · Contributes. Source · Evaluator-reported · Muse Spark 1.2 (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 45.458758% · Contributes. Source · Evaluator-reported · Muse Spark 1.2 (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 43.75% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 25.42% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1346.23 · Contributes. Source · Evaluator-reported · Muse Spark 1.2 (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 90.06% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 86.6% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 85.26% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 80.38% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- KingBench 3: 76.25% · Not used.
Not admitted after the September 6 model review: the collected 76.25 for Muse Spark 1.2 on KingBench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Index: 71.88% · Not used. Source · Evaluator-reported · muse-spark-1.2; xhigh effort; Vals AI index methodology
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 68.47% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 49.35% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AutomationBench v1.0.6: 38.2% · Not used.
Not admitted after the September 6 model review: the collected 38.2 for Muse Spark 1.2 on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 27.2 · Contributes. Source · Evaluator-reported · Muse Spark 1.2 (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 88.28% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 78% · Not used.
Not admitted after the September 6 model review: the collected 78 for Muse Spark 1.2 on LiveBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CorpFin v2 ; Vals: 70.94% · Contributes. Source · Archived source review · muse-spark-1.2; xhigh effort; Vals AI CorpFin v2 evaluation with Sonnet 4.5 judge
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MortgageTax ; Vals: 65.42% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 47.66% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 29.95% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Toolathlon-Verified: 75.9% · Contributes. Source · Evaluator-reported · Meta Muse Spark 1.2 (xhigh) ✓; Default agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Design Arena (Elo): 1373 · Contributes. Source · Archived source review · muse-spark-1.2; xhigh effort; Design Arena 2026-08-11 rating snapshot
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Vals Multimodal Index: 69.8% · Not used. Source · Evaluator-reported · muse-spark-1.2; xhigh effort; Vals AI multimodal index methodology
Held-out composite, never a scoring input
- SkillsBench †: 53.04% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- OSWorld 2.0 ; Anthropic H2H: 47.6% · Not used.
Not admitted after the September 6 model review: the collected 47.6 for Muse Spark 1.2 on OSWorld 2.0 ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- JobBench: 61.6% · Not used.
Not admitted after the September 6 model review: the collected 61.6 for Muse Spark 1.2 on JobBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Creative Writing v3: 1835.4 · Not used.
Not admitted after the September 6 model review: the collected 1835.4 for Muse Spark 1.2 on Creative Writing v3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MCP Atlas †: 90.3% · Not used.
Not admitted after the September 6 model review: the collected 90.3 for Muse Spark 1.2 on MCP Atlas † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 34.845361% · Contributes. Source · Evaluator-reported · Muse Spark 1.2 (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 17.714286% · Contributes. Source · Evaluator-reported · Muse Spark 1.2 (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- GDP.pdf: 16% · Contributes. Source · Evaluator-reported · Muse Spark 1.2 (xHigh reasoning); Surge original evaluation
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 43% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 79.1% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 1653 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- OSWorld 2.0: 17.9% · Not used.
Not admitted after the September 6 model review: the collected 17.9 for Muse Spark 1.2 on OSWorld 2.0 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EnterpriseOps-Gym-AA: 47.27% · Not used.
Not admitted after the September 6 model review: the collected 47.27 for Muse Spark 1.2 on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 31.111111% · Contributes. Source · Evaluator-reported · Muse Spark 1.2 (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CWE-bench: 30.3% · Not used.
Not admitted after the September 6 model review: the collected 30.3 for Muse Spark 1.2 on CWE-bench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- BioMysteryBench v1 ; Vals: 64.81% · Contributes. Source · Evaluator-reported · meta_muse_spark_1_2; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- RuneBench (30m, mean ln(1 + XP/min)): 5.238233 · Contributes. Source · Archived source review · muse12
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Roboflow Vision Evals (six-task mean, high tier): 80.016667% · Contributes. Source · Archived source review · Require a complete six-task high-tier panel, the highest published benchmark tier. Retain high-tier scores even when lower than low-tier. Do not substitute the best of three repetitions: use the published mean. Each panel enters the fit once; do not also fit its components.
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1534 · Contributes. Source · Archived source review · muse-spark-1.2 (xHigh)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
DeepSeek V4 Pro · 38 used / 56 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
36 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.4 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1496.68 · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 78.651685% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 62.831858% · Contributes. Source · Evaluator-reported · deepseek-v4-pro; max; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1459.610748 · Contributes. Source · Evaluator-reported · deepseek-v4-pro-high-20260813; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 53 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 92.828283% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 75.3% · Not used. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 50.39% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 52.8% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 51.041667% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 41.010195% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 40.87% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 7.5% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1269.34 · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 80.17% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 96.4% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 82.36% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 73.06% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- KingBench 3: 76.25% · Not used.
Not admitted after the September 6 model review: the collected 76.25 for DeepSeek V4 Pro Max on KingBench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Index: 66.25% · Not used. Source · Evaluator-reported · deepseek-v4-pro; max effort; Vals AI index methodology
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 62.92% · Contributes. Source · Archived source review · deepseek-v4-pro; max effort; Vals AI public-benefits evaluation
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- MedCode ; Vals: 42.47% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AutomationBench v1.0.6: 43.2% · Not used.
Not admitted after the September 6 model review: the collected 43.2 for DeepSeek V4 Pro Max on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: 0.833333 · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 86.97% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 77.436143% · Contributes. Source · Evaluator-reported · deepseek-v4-pro-0813; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- CorpFin v2 ; Vals: 65.42% · Contributes. Source · Archived source review · deepseek-v4-pro; max effort; Vals AI CorpFin v2 evaluation with Sonnet 4.5 judge
Exact-value match to an existing verified, comparable, pinned-configuration observation. Archived review retained; not asserted newly independently replicated.
- Code Migration †: 41.54% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Toolathlon-Verified: 74.4% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (max) ✓; Default agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Agents' Last Exam: 12.4% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro Max; High; OpenClaw
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveCodeBench: 87.53% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CyberGym: 83.3% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro Max 0813; official dated DeepSeek release aggregate
DeepSeek August 13 (Pro 0813) / July 31 (Flash 0731) official release. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- Humanity's Last Exam w/ tools: 60% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro Max 0813; official dated DeepSeek release aggregate
DeepSeek August 13 (Pro 0813) / July 31 (Flash 0731) official release. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- SWE-Marathon v1.1: 10.6% · Not used.
Not admitted after the September 6 model review: the collected 10.6 for DeepSeek V4 Pro Max on SWE-Marathon v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 70.103% · Contributes. Source · Archived source review · deepseek/deepseek-v4-pro-0813 · reasoning_effort=max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- FrontierMath v2 Tier 4: 26.83% · Not used.
Not admitted after the September 6 model review: the collected 26.83 for DeepSeek V4 Pro Max on FrontierMath v2 Tier 4 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SWE-bench Pro: 55.4% · Not used.
Not admitted after the September 6 model review: the collected 55.4 for DeepSeek V4 Pro Max on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- BrowseComp: 83.4% · Not used.
Not admitted after the September 6 model review: the collected 83.4 for DeepSeek V4 Pro Max on BrowseComp still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ARC-AGI-2: 61.3% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro Max; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- SimpleQA: 57.9% · Not used.
Not admitted after the September 6 model review: the collected 57.9 for DeepSeek V4 Pro Max on SimpleQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MCP Atlas †: 73.6% · Not used.
Not admitted after the September 6 model review: the collected 73.6 for DeepSeek V4 Pro Max on MCP Atlas † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- NL2Repo: 61.5% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro Max 0813; official dated DeepSeek release aggregate
DeepSeek August 13 (Pro 0813) / July 31 (Flash 0731) official release. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial. Corrects the formerly unresolved 61.1 to the model owner’s explicit 0813 result, 61.5; the old observation remains archived.
- CoWorkBench: 66.3% · Not used.
Not admitted after the September 6 model review: the collected 66.3 for DeepSeek V4 Pro Max on CoWorkBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ²-Bench Telecom: 96.2% · Not used.
Not admitted after the September 6 model review: the collected 96.2 for DeepSeek V4 Pro Max on τ²-Bench Telecom still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Main ; Anthropic H2H: 17.6% · Not used.
Not admitted after the September 6 model review: the collected 17.6 for DeepSeek V4 Pro Max on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EQ-Bench 3: 1570.2 · Not used.
Not admitted after the September 6 model review: the collected 1570.2 for DeepSeek V4 Pro Max on EQ-Bench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 39.587629% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 18% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ProofBench v1.1 ; Vals: 50% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 82.3% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ARC-AGI-1: 90.5% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro Max; Low; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- EnterpriseOps-Gym-AA: 49.6% · Not used.
Not admitted after the September 6 model review: the collected 49.6 for DeepSeek V4 Pro Max on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 17.777778% · Contributes. Source · Evaluator-reported · DeepSeek V4 Pro 0813 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- IFBench: 76.5% · Not used.
Not admitted after the September 6 model review: the collected 76.5 for DeepSeek V4 Pro Max on IFBench still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Mystery Game Puzzles v1.0.4 ; Epoch: 43% · Not used.
Not admitted after the September 6 model review: the collected 43 for DeepSeek V4 Pro Max on Mystery Game Puzzles v1.0.4 ; Epoch still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Tax Agent Bench v1 ; Vals: 58.66% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-pro-0813; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
GPT-5.6 Luna · 54 used / 77 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
50 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.2 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1492.73 · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 80.898876% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 67.1875% · Contributes. Source · Evaluator-reported · gpt-5-6-luna; max; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1452.599383 · Contributes. Source · Evaluator-reported · gpt-5.6-luna-xhigh; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 52 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 91.111111% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 78.3% · Not used. Source · Evaluator-reported · GPT-5.6 Luna (max)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 55.04% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 67.12% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 53.587963% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 39.481001% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 36.54% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 1.25% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1342.42 · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 84.39% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 93% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 84.03% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 76.17% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vals Index: 59.88% · Not used.
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 61.16% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 42.39% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Omniscience: -10.283333 · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 86.04% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 73.559381% · Contributes. Source · Evaluator-reported · gpt-5.6-luna-max; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- CorpFin v2 ; Vals: 64.22% · Not used.
Not admitted after the September 6 model review: the collected 64.22 for GPT-5.6 Luna on CorpFin v2 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MortgageTax ; Vals: 67.29% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 44.22% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 44.55% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Agents' Last Exam: 30.3% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna; XHigh; Codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Terminal-Bench v3.0: 14.3% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); Codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CyberGym: 77.9% · Not used.
Not admitted after the September 6 model review: the collected 77.9 for GPT-5.6 Luna on CyberGym still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Humanity's Last Exam w/ tools: 48.9% · Not used.
Not admitted after the September 6 model review: the collected 48.9 for GPT-5.6 Luna on Humanity's Last Exam w/ tools still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Design Arena (Elo): 1283 · Not used.
Not admitted after the September 6 model review: the collected 1283 for GPT-5.6 Luna on Design Arena (Elo) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Multimodal Index: 69.06% · Not used.
Held-out composite, never a scoring input
- SWE-Marathon v1.1: 24.4% · Not used.
Not admitted after the September 6 model review: the collected 24.4 for GPT-5.6 Luna on SWE-Marathon v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 68.291% · Contributes. Source · Archived source review · openai/gpt-5.6-luna · reasoning_effort=max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- SkillsBench †: 60.45% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierMath v2 Tier 4: 60.97561% · Contributes. Source · Evaluator-reported · gpt-5.6-luna_max; task version 2.0.0
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- CursorBench v3.2: 61.1% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna Max; Cursor harness
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Pro: 62.7% · Not used.
Not admitted after the September 6 model review: the collected 62.7 for GPT-5.6 Luna on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- BrowseComp: 83.3% · Not used.
Not admitted after the September 6 model review: the collected 83.3 for GPT-5.6 Luna on BrowseComp still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MMMU-Pro: 78.554913% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (xhigh); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ARC-AGI-2: 59.5% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- OSWorld 2.0 ; Anthropic H2H: 45.6% · Not used.
Not admitted after the September 6 model review: the collected 45.6 for GPT-5.6 Luna on OSWorld 2.0 ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ExploitBench: 33.2% · Not used.
Inherited comparator metric/scaffold equivalence to Astra Cap Percent is not established; the source flags possible contamination. Excluded pending protocol audit.
- SimpleQA: 41.7% · Not used.
Not admitted after the September 6 model review: the collected 41.7 for GPT-5.6 Luna on SimpleQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- FrontierCode v1.1 Extended: 55.06% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna; max; codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Creative Writing v3: 1826.6 · Contributes. Source · Evaluator-reported · gpt-5.6-luna; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- AIIQ Composite IQ: 129 · Not used.
Held-out composite, never a scoring input
- EQ-Bench 4: 1156.3 · Contributes. Source · Evaluator-reported · openai/gpt-5.6-luna; evaluator published configuration, effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ERQA: 62.42% · Not used.
Not admitted after the September 6 model review: the collected 62.42 for GPT-5.6 Luna on ERQA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AutomationBench ; Anthropic H2H: 14.9% · Not used.
Not admitted after the September 6 model review: the collected 14.9 for GPT-5.6 Luna on AutomationBench ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 31.134021% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 20.571429% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (xhigh); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 4.0: 17.3% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna; max; Codex
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- GDP.pdf: 22.7% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (Medium reasoning); Surge original evaluation
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 60% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 77.06% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 1784 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- ARC-AGI-1: 88% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- ARC-AGI-3 (standard harness): 0.18% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- HealthBench: 55.4% · Not used. Source · Developer-reported · GPT-5.6 Luna; OpenAI published HealthBench research evaluation; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- HealthBench Hard: 31.4% · Not used. Source · Developer-reported · GPT-5.6 Luna; OpenAI published HealthBench research evaluation; effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- HealthBench Professional (length-adjusted): 55.7% · Not used.
Source selects the maximum score across efforts; exact highest-effort result is not established.
- Harvey LAB-AA: 87.9% · Not used.
Not admitted after the September 6 model review: the collected 87.9 for GPT-5.6 Luna on Harvey LAB-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EnterpriseOps-Gym-AA: 40.82% · Not used.
Not admitted after the September 6 model review: the collected 40.82 for GPT-5.6 Luna on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 19.444444% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ITBench-AA: 40.320151% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Vending-Bench 2: 4094.712 · Contributes. Source · Evaluator-reported · GPT-5.6 Luna; Andon Labs Vending-Bench 2
Compare reasoning configurations using their mean over runs. Geometric mean, individual best runs, Fireworks Kimi endpoint, Vending Arena, and April DeepSeek Pro are not substituted.
- ZeroBench: 21% · Contributes. Source · Evaluator-reported · GPT-5.6 Luna (max); official main questions
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Mystery Game Puzzles v1.0.4 ; Epoch: 21% · Contributes. Source · Evaluator-reported · gpt-5.6-luna_max; task version 1.0.4
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- BioMysteryBench v1 ; Vals: 61.48% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Tax Agent Bench v1 ; Vals: 60.81% · Contributes. Source · Evaluator-reported · openai_gpt-5.6-luna; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- RuneBench (30m, mean ln(1 + XP/min)): 5.285493 · Contributes. Source · Archived source review · gpt56luna-xhigh
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 1846 · Contributes. Source · Archived source review · GPT-5.6-Luna
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Roboflow Vision Evals (six-task mean, high tier): 76.05% · Contributes. Source · Archived source review · Require a complete six-task high-tier panel, the highest published benchmark tier. Retain high-tier scores even when lower than low-tier. Do not substitute the best of three repetitions: use the published mean. Each panel enters the fit once; do not also fit its components.
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1519 · Contributes. Source · Archived source review · gpt-5.6-luna-xhigh (codex-harness)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
Sonnet 5 · 51 used / 70 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
49 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.2 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1504.87 · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 6 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 80.524345% · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 53.846154% · Contributes. Source · Evaluator-reported · claude-sonnet-5; max; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1462.365844 · Contributes. Source · Evaluator-reported · claude-sonnet-5-high; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 55 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 91.111111% · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 77% · Not used. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 53.91% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 66.32% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 54.282407% · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 41.28823% · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 41.83% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 5% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1358.11 · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 5 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 76.05% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 79.6% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 83.92% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 75.63% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vals Index: 59.61% · Not used.
Held-out composite, never a scoring input
- Public Benefits Bench ; Vals: 66.03% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 47.54% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Omniscience: 16.45 · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 87.55% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 76.039714% · Contributes. Source · Evaluator-reported · claude-sonnet-5-xhigh-effort; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- CorpFin v2 ; Vals: 67.95% · Not used.
Not admitted after the September 6 model review: the collected 67.95 for Sonnet 5 on CorpFin v2 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MortgageTax ; Vals: 70.03% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 48.92% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 44.39% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Toolathlon-Verified: 71.6% · Contributes. Source · Evaluator-reported · Claude Claude Sonnet 5 (max) ✓; Default agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Terminal-Bench v3.0: 14.6% · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (max); Claude Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveCodeBench: 82.43% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CyberGym: 52.7% · Not used.
Not admitted after the September 6 model review: the collected 52.7 for Sonnet 5 on CyberGym still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Humanity's Last Exam w/ tools: 57.4% · Not used.
Not admitted after the September 6 model review: the collected 57.4 for Sonnet 5 on Humanity's Last Exam w/ tools still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Design Arena (Elo): 1297 · Not used.
Not admitted after the September 6 model review: the collected 1297 for Sonnet 5 on Design Arena (Elo) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Multimodal Index: 68.83% · Not used.
Held-out composite, never a scoring input
- ProgramBench v1 (Raw Pass Rate) ; Vals: 72.067% · Contributes. Source · Archived source review · anthropic/claude-sonnet-5 · compute=max
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- SkillsBench †: 46.48% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierMath v2 Tier 4: 29.268293% · Contributes. Source · Evaluator-reported · claude-sonnet-5_max; task version 2.0.0
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- CursorBench v3.2: 61.5% · Contributes. Source · Evaluator-reported · Sonnet 5 Max; Cursor harness
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- APEX-Agents Mean Criteria Passed: 48.5% · Contributes. Source · Evaluator-reported · claude-sonnet-5; high; loop_truncated_tools_agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Pro: 63.2% · Not used.
Not admitted after the September 6 model review: the collected 63.2 for Sonnet 5 on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- CharXiv (with CI / RQ): 88.3% · Not used.
Not admitted after the September 6 model review: the collected 88.3 for Sonnet 5 on CharXiv (with CI / RQ) still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- APEX-SWE (Pass@1, Terminus-2): 46.4% · Contributes. Source · Evaluator-reported · claude-sonnet-5-max; max; terminus-2
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- BrowseComp: 84.7% · Not used.
Not admitted after the September 6 model review: the collected 84.7 for Sonnet 5 on BrowseComp still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MMMU-Pro: 77.283237% · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- FrontierCode v1.1 Extended: 56.18% · Contributes. Source · Evaluator-reported · Claude Sonnet 5; xhigh; claude-code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- OSWorld-Verified: 81.2% · Not used.
Not admitted after the September 6 model review: the collected 81.2 for Sonnet 5 on OSWorld-Verified still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Creative Writing v3: 1787.6 · Contributes. Source · Evaluator-reported · claude-sonnet-5; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- OfficeQA Pro †: 59.4% · Not used.
Not admitted after the September 6 model review: the collected 59.4 for Sonnet 5 on OfficeQA Pro † still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EQ-Bench 4: 1236 · Contributes. Source · Evaluator-reported · claude-sonnet-5; evaluator published configuration, effort unspecified
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierCode v1.1 Main ; Anthropic H2H: 42.7% · Not used.
Not admitted after the September 6 model review: the collected 42.7 for Sonnet 5 on FrontierCode v1.1 Main ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AutomationBench ; Anthropic H2H: 13.5% · Not used.
Not admitted after the September 6 model review: the collected 13.5 for Sonnet 5 on AutomationBench ; Anthropic H2H still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Harvey Legal Agent Benchmark ; held-out: 5.8% · Not used.
Not admitted after the September 6 model review: the collected 5.8 for Sonnet 5 on Harvey Legal Agent Benchmark ; held-out still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 37.319588% · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 16.857143% · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 4.0: 12.4% · Contributes. Source · Evaluator-reported · Sonnet 5; max; Claude Code
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProofBench v1.1 ; Vals: 77% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 81.33% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 1720 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- Harvey LAB-AA: 90.07% · Not used.
Not admitted after the September 6 model review: the collected 90.07 for Sonnet 5 on Harvey LAB-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EnterpriseOps-Gym-AA: 44.67% · Not used.
Not admitted after the September 6 model review: the collected 44.67 for Sonnet 5 on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 55% · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-AnalystAgent: 46.25% · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (Adaptive Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Vending-Bench 2: 6377.701667 · Contributes. Source · Evaluator-reported · Claude Sonnet 5; Andon Labs Vending-Bench 2
Compare reasoning configurations using their mean over runs. Geometric mean, individual best runs, Fireworks Kimi endpoint, Vending Arena, and April DeepSeek Pro are not substituted.
- ZeroBench: 13% · Contributes. Source · Evaluator-reported · Claude Sonnet 5 (max); official main questions
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Mystery Game Puzzles v1.0.4 ; Epoch: 35% · Contributes. Source · Evaluator-reported · claude-sonnet-5_max; task version 1.0.4
Exact Epoch task version; selected best mean score across reported reasoning configurations. No pooling of task versions or substitution of best scorer output.
- Tax Agent Bench v1 ; Vals: 62.27% · Contributes. Source · Evaluator-reported · anthropic_claude-sonnet-5; Vals max effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- RuneBench (30m, mean ln(1 + XP/min)): 4.805888 · Contributes. Source · Archived source review · sonnet5-xhigh
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 1570 · Contributes. Source · Archived source review · Claude Sonnet 5 (xhigh)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1537 · Contributes. Source · Archived source review · claude-sonnet-5-high
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
Qwen3.8-27B · 40 used / 52 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
40 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.3 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1466.21 · Contributes. Source · Evaluator-reported · Qwen3.8 27B (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 79.775281% · Contributes. Source · Evaluator-reported · Qwen3.8 27B (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 42.2% · Not used.
Not admitted after the September 6 model review: the collected 42.2 for Qwen3.8-27B on DeepSWE v1.1 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- LMArena (Chatbot Arena) ; Text: 1435.892663 · Contributes. Source · Evaluator-reported · qwen3.8-27b; Arena text overall; style control enabled
Best published reasoning-mode rating within the same text overall board. Different leaderboard slices and undated earlier DeepSeek checkpoints are not pooled.
- AA Intelligence Index v4.1.1 (score): 52 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 90.505051% · Contributes. Source · Evaluator-reported · Qwen3.8 27B (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 77.3% · Not used. Source · Evaluator-reported · Qwen3.8 27B (xhigh)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 48.55% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 59.66% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 46.643519% · Contributes. Source · Evaluator-reported · Qwen3.8 27B (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 33.920297% · Contributes. Source · Evaluator-reported · Qwen3.8 27B (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 36.06% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 11.25% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1401.81 · Contributes. Source · Evaluator-reported · Qwen3.8 27B (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 83.85% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 86% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 82.43% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 70.85% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MedCode ; Vals: 28.7% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Omniscience: -7.95 · Contributes. Source · Evaluator-reported · Qwen3.8 27B (Non-reasoning); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 84.34% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 75.269107% · Contributes. Source · Evaluator-reported · qwen3.8-27b; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- MortgageTax ; Vals: 64.94% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SAGE ; Vals: 52.4% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Code Migration †: 14.16% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Agents' Last Exam: 20.4% · Not used.
Not admitted after the September 6 model review: the collected 20.4 for Qwen3.8-27B on Agents' Last Exam still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- LiveCodeBench: 84% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ProgramBench v1 (Raw Pass Rate) ; Vals: 11.168% · Contributes. Source · Archived source review · alibaba/qwen3.8-27b · reasoning_effort=xhigh
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- SWE-bench Pro: 61.7% · Not used. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CharXiv (with CI / RQ): 90.2% · Contributes. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; With CI
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MMMU-Pro: 76.300578% · Contributes. Source · Evaluator-reported · Qwen3.8 27B (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- JobBench: 33.4% · Contributes. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- BabyVision (with CI): 85.6% · Contributes. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; With CI
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- OSWorld-Verified: 84.3% · Contributes. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Creative Writing v3: 1669.1 · Contributes. Source · Evaluator-reported · *Qwen/Qwen3.8-27B; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- NL2Repo: 42.3% · Contributes. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CoWorkBench: 70.7% · Not used. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- ERQA: 65.5% · Not used. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vision2Web (Avg. Frontend/Webpage/etc.): 62.9% · Not used. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- WebArena-Verified: 64.8% · Not used. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- τ³-Banking: 48.041237% · Contributes. Source · Evaluator-reported · Qwen3.8 27B (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 5.428571% · Contributes. Source · Evaluator-reported · Qwen3.8 27B (xhigh); Artificial Analysis published evaluation harness
Best published result among 4 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ProofBench v1.1 ; Vals: 16% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 64.85% · Contributes. Source · Evaluator-reported · alibaba_qwen3.8-27b; Vals xhigh effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- EnterpriseOps-Gym-AA: 44.23% · Not used.
Not admitted after the September 6 model review: the collected 44.23 for Qwen3.8-27B on EnterpriseOps-Gym-AA still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MLCR-AA (Medical Long Context Reasoning): 21.666667% · Contributes. Source · Evaluator-reported · Qwen3.8 27B (xhigh); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- IFBench: 79.5% · Not used. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AndroidWorld: 81.9% · Not used. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MathVision (with CI): 94.6% · Contributes. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; With CI
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- OmniDocBench 1.5: 91.1% · Contributes. Source · Developer-reported · Qwen3.8-27B; published developer evaluation; benchmark-specific configuration
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- RuneBench (30m, mean ln(1 + XP/min)): 2.794282 · Contributes. Source · Archived source review · qwen38
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- Code Arena WebDev (overall): 1594 · Contributes. Source · Archived source review · qwen3.8-27b
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
DeepSeek V4 Flash · 39 used / 57 collected
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
38 distinct benchmark families contribute. Domains without direct evidence are estimates.
Reporting sensitivity: ±1.3 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
- GDPval-AA v2 (Elo): 1580.66 · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash Vision (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Terminal-Bench 2.1: 78.651685% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash 0731 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- DeepSWE v1.1: 53.318584% · Contributes. Source · Evaluator-reported · deepseek-v4-flash; max; mini-swe-agent
Selected the best pass@1 across published reasoning efforts within mini-swe-agent. Pass@4 is not substituted.
- LMArena (Chatbot Arena) ; Text: 1435 · Not used.
Not admitted after the September 6 model review: the collected 1435 for DeepSeek V4 Flash on LMArena (Chatbot Arena) ; Text still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA Intelligence Index v4.1.1 (score): 52 · Not used.
Held-out composite, never a scoring input
- GPQA Diamond: 91.313131% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash Vision (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- AA-LCR: 66% · Not used. Source · Evaluator-reported · DeepSeek V4 Flash 0731 (Reasoning, Max Effort)
Current AA result is LCR v1.1. The inherited AA-LCR row has no reviewed version match; a new version cannot silently substantiate the older value.
- Finance Agent v2 ; Vals: 49.52% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Excel Modeling Benchmark ; Vals: 56.98% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SciCode: 50.347222% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash 0731 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Humanity's Last Exam (no tools): 38.554217% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash 0731 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- Legal Research Bench ; Vals: 30.29% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Harvey LAB (Vals): 8.33% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AA-Briefcase Overall Elo: 1262.58 · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash 0731 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MedScribe ; Vals: 80.36% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Verified ; Vals: 88.8% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LegalBench ; Vals: 77.71% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- TaxEval v2 ; Vals: 70.69% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- KingBench 3: 72.5% · Not used.
Not admitted after the September 6 model review: the collected 72.5 for DeepSeek V4 Flash on KingBench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Vals Index: 53.57% · Not used.
Held-out composite, never a scoring input
- MedCode ; Vals: 41.41% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- AutomationBench v1.0.6: 25.1% · Not used.
Not admitted after the September 6 model review: the collected 25.1 for DeepSeek V4 Flash on AutomationBench v1.0.6 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- AA-Omniscience: -14.283333 · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash 0731 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- MMLU Pro ; Vals: 86.21% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- LiveBench: 74.170786% · Contributes. Source · Evaluator-reported · deepseek-v4-flash-0731; LiveBench 2026-06-25 release
Task mapping archived from categories_2026_06_25.json. Aggregate follows the official UI, not a flat mean of the 23 tasks. Undated old DeepSeek checkpoints excluded.
- CorpFin v2 ; Vals: 61.85% · Not used.
Not admitted after the September 6 model review: the collected 61.85 for DeepSeek V4 Flash on CorpFin v2 ; Vals still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- Code Migration †: 38.63% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Toolathlon-Verified: 70.7% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash 0731 (max) ✓; Default agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Agents' Last Exam: 25.2% · Not used.
Not admitted after the September 6 model review: the collected 25.2 for DeepSeek V4 Flash on Agents' Last Exam still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- LiveCodeBench: 87.26% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- CyberGym: 76.7% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash 0731; official dated DeepSeek release aggregate
DeepSeek August 13 (Pro 0813) / July 31 (Flash 0731) official release. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- Humanity's Last Exam w/ tools: 45.1% · Not used.
Not admitted after the September 6 model review: the collected 45.1 for DeepSeek V4 Flash on Humanity's Last Exam w/ tools still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- SkillsBench †: 50.67% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- FrontierMath v2 Tier 4: 24.4% · Not used.
Not admitted after the September 6 model review: the collected 24.4 for DeepSeek V4 Flash on FrontierMath v2 Tier 4 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- APEX-Agents Mean Criteria Passed: 51.6% · Contributes. Source · Evaluator-reported · deepseek-v4-flash; max; loop_truncated_tools_agent
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- SWE-bench Pro: 52.6% · Not used.
Not admitted after the September 6 model review: the collected 52.6 for DeepSeek V4 Flash on SWE-bench Pro still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- APEX-SWE (Pass@1, Terminus-2): 46.9% · Contributes. Source · Evaluator-reported · deepseek-v4-flash; max; terminus-2
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- BrowseComp: 73.2% · Not used.
Not admitted after the September 6 model review: the collected 73.2 for DeepSeek V4 Flash on BrowseComp still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- ARC-AGI-2: 61.4% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- FrontierCode v1.1 Extended: 31.7% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash 0731; high; chisel
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Creative Writing v3: 1438.3 · Contributes. Source · Evaluator-reported · *deepseek-ai/DeepSeek-V4-Flash-0731; published evaluation, effort unspecified
Exact checkpoint on the evaluator-hosted board. A starred model is externally submitted; it is not claimed as independently rerun by the benchmark owner.
- MathArena ; ArXivMath Jun 2026: 42.86% · Not used.
Not admitted after the September 6 model review: the collected 42.86 for DeepSeek V4 Flash on MathArena ; ArXivMath Jun 2026 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- MathArena ; BrokenArXiv Jun 2026: 16.67% · Not used.
Not admitted after the September 6 model review: the collected 16.67 for DeepSeek V4 Flash on MathArena ; BrokenArXiv Jun 2026 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- NL2Repo: 54.2% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash 0731; official dated DeepSeek release aggregate
DeepSeek August 13 (Pro 0813) / July 31 (Flash 0731) official release. Reviewed developer-reported result for the exact named reference model. Cross-evaluator harness differences are retained as a limitation; no independent evaluation or matched-harness claim. Sole reported aggregate, not a selected individual trial.
- MathArena ; ArXivMath May 2026: 55.83% · Not used.
Not admitted after the September 6 model review: the collected 55.83 for DeepSeek V4 Flash on MathArena ; ArXivMath May 2026 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- EQ-Bench 3: 1491.3 · Not used.
Not admitted after the September 6 model review: the collected 1491.3 for DeepSeek V4 Flash on EQ-Bench 3 still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- τ³-Banking: 41.030928% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash Vision (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CritPt: 16.571429% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash 0731 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 2 reasoning configurations of the exact AA model release; fallback configurations excluded.
- ProofBench v1.1 ; Vals: 56% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Vibe Code Bench v1.1 ; Vals: 74.74% · Contributes. Source · Evaluator-reported · deepseek_deepseek-v4-flash-0731; Vals high effort
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- MineBench: 1601 · Not used.
Unresolved Pro/base model identities in source catalog; exclude until matched.
- ARC-AGI-1: 89% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash; Max; ARC Prize standard harness
Best reasoning variant within the verified standard-harness summary. Provider Adapter, public tasks and community agent scores are not substituted.
- MLCR-AA (Medical Long Context Reasoning): 13.333333% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash 0731 (Reasoning, Max Effort); Artificial Analysis published evaluation harness
Best published result among 1 reasoning configurations of the exact AA model release; fallback configurations excluded.
- CWE-bench: 30.4% · Contributes. Source · Evaluator-reported · DeepSeek V4 Flash; high; opencode
Reconciled against the exact model/configuration in the evaluator’s public source. Updated source value recorded explicitly; best published reasoning mode selected within the same benchmark protocol.
- Mystery Game Puzzles v1.0.4 ; Epoch: 34% · Not used.
Not admitted after the September 6 model review: the collected 34 for DeepSeek V4 Flash on Mystery Game Puzzles v1.0.4 ; Epoch still lacks a matched source entry establishing its checkpoint and benchmark-specific metric, tools and retry setup. This is a citation/comparability gap, not a finding that the result is false or useless. More or maximum reasoning effort is not required.
- RuneBench (30m, mean ln(1 + XP/min)): 4.195298 · Contributes. Source · Archived source review · deepseekflash0731
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.
- VoxelBench (text-to-voxel, Glicko-2): 1436 · Contributes. Source · Archived source review · DeepSeek V4 Flash (July 2026)
Matched exact metric/value and the archived benchmark-specific configuration review. Unknown native effort stays explicitly unknown.