Gemini 4 Argon: score and benchmark evidence

Historical release · · Google

134.8 points · Rank 3 among 31 ranked models · Stability range 123.4–146.9 points.

Rank follows the point estimate; it does not establish statistically significant superiority. The scale is anchored at mean 100, SD 15 in the frozen calibration cohort, not human IQ or a percentage.

One score summarizes demonstrated capability across published reasoning settings and seven equally weighted domains. Related benchmark variants share a family budget, and influence from one evaluator is capped.

Release tli-2026-v1.0-2026-10-10-refresh-21. Methodology and limitations for this release.

This page preserves historical evidence. See the current model record.

Capability domains

Domain scores use calibration-cohort standard-deviation units. Unmeasured domains are shown as unavailable, not zero.

DomainScore (SD units)Observed benchmarks
Knowledge & reasoning1.022
Coding & software engineering1.214
Agentic tool & computer use0.801
Professional & real-world work1.122
Multimodal & visionUnavailable0
CybersecurityUnavailable0
Preference & communicationUnavailable0

Published benchmark evidence

9 results contribute to this score, from 17 collected results and 9 contributing benchmark families. Counts are not independent sample sizes.

Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.

The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.

Reporting sensitivity: ±21.6 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.

  1. AA-Briefcase v1.1 (Elo): 1488.47 native-points

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Gemini 4 Argon (high); Artificial Analysis independent evaluation, 2026-10-10; fallback not reported

    Metric: published native-points

    Original source · Retrieved 2026-10-10

    Daily evaluator refresh; replaces 1489.88 recorded 2026-10-06.

  2. GDPval-AA v2.1 (Elo): 1623.83 native-points

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Gemini 4 Argon (high); Artificial Analysis independent evaluation, 2026-10-10; fallback not reported

    Metric: published native-points

    Original source · Retrieved 2026-10-10

    Daily evaluator refresh; replaces 1626.19 recorded 2026-10-06.

  3. AutomationBench-AA: 77.513469459%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Gemini 4 Argon (high); Artificial Analysis independent evaluation, 2026-09-30; fallback not reported

    Metric: published percent

    Original source · Retrieved 2026-09-30

    Exact model and effort in the evaluator chart payload.

  4. Terminal-Bench 4.0 — AA harness: 57.0707070707%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Gemini 4 Argon (high); Artificial Analysis independent evaluation, 2026-09-30; fallback not reported

    Metric: published percent

    Original source · Retrieved 2026-09-30

    Exact model and effort in the evaluator chart payload.

  5. SciCode: 61.8055555556%

    Contributes to the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Gemini 4 Argon (high); Artificial Analysis independent evaluation, 2026-09-30; fallback not reported

    Metric: published percent

    Original source · Retrieved 2026-09-30

    Exact model and effort in the evaluator chart payload.

  6. HLE text-only — AA harness: 57.0898980538%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Gemini 4 Argon (high); Artificial Analysis independent evaluation, 2026-09-30; fallback not reported

    Metric: published percent

    Original source · Retrieved 2026-09-30

    Exact model and effort in the evaluator chart payload.

  7. GDP.pdf — AA harness: 21.8%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Gemini 4 Argon (high); Artificial Analysis independent evaluation, 2026-09-30; fallback not reported

    Metric: published percent

    Original source · Retrieved 2026-09-30

    Exact model and effort in the evaluator chart payload.

  8. CritPt: 27.1428571429%

    Contributes to the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Gemini 4 Argon (high); Artificial Analysis independent evaluation, 2026-09-30; fallback not reported

    Metric: published percent

    Original source · Retrieved 2026-09-30

    Exact model and effort in the evaluator chart payload.

  9. AA-Omniscience: 42.35 native-points

    Contributes to the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Gemini 4 Argon (high); Artificial Analysis independent evaluation, 2026-09-30; fallback not reported

    Metric: published native-points

    Original source · Retrieved 2026-09-30

    Exact model and effort in the evaluator chart payload.

  10. AA-LCR v1.1: 79.6666666667%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Gemini 4 Argon (high); Artificial Analysis independent evaluation, 2026-09-30; fallback not reported

    Metric: published percent

    Original source · Retrieved 2026-09-30

    Exact model and effort in the evaluator chart payload.

  11. AA Intelligence Index v4.3.2 (score): 52.5605655746 native-points

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Gemini 4 Argon (high); Artificial Analysis independent evaluation, 2026-09-30; fallback not reported

    Metric: published native-points

    Original source · Retrieved 2026-09-30

    Reference composite only, never scored alongside its component evaluations.

  12. DeepSWE v1.1: 77.9%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Gemini 4 Argon (effort not stated); Google Gemini 4 Argon launch benchmark table; methodology at deepmind.google/models/evals-methodology/gemini-4-argon; fallback not reported

    Metric: published score (%)

    Original source · Retrieved 2026-10-08

    Developer-reported aggregate. Same protocol as the catalog row: the lab table reproduces existing cells (gpt-6-astra 74.1, opus-5-5 74.2, fable-5-1 67.4). Effort not stated.

  13. Vibe Code Bench v1.1 — Vals: 91.9%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Gemini 4 Argon (effort not stated); Google Gemini 4 Argon launch benchmark table; methodology at deepmind.google/models/evals-methodology/gemini-4-argon; fallback not reported

    Metric: published score (%)

    Original source · Retrieved 2026-10-08

    Developer-reported aggregate. Same protocol as the catalog row: the lab table reproduces existing cells (gpt-6-astra 89.6, fable-5-1 90.3). Effort not stated. Table does not name the version; comparison columns match v1.1.

  14. Terminal-Bench 4.0: 57.4%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Gemini 4 Argon (effort not stated); Google Gemini 4 Argon launch benchmark table; methodology at deepmind.google/models/evals-methodology/gemini-4-argon; fallback not reported

    Metric: published score (%)

    Original source · Retrieved 2026-10-08

    Developer-reported aggregate. Same protocol as the catalog row: the lab table reproduces existing cells (gpt-6-astra 58.2, fable-5-1 57.9, opus-5-5 66.4). Effort not stated.

  15. Finance Agent v2 — Vals: 65.4%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Gemini 4 Argon (effort not stated); Google Gemini 4 Argon launch benchmark table; methodology at deepmind.google/models/evals-methodology/gemini-4-argon; fallback not reported

    Metric: published score (%)

    Original source · Retrieved 2026-10-08

    Developer-reported aggregate. Same protocol as the catalog row: the lab table reproduces existing cells (gpt-6-astra 53.5, fable-5-1 58.9). Effort not stated.

  16. Harvey LAB (Vals): 19.6%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Gemini 4 Argon (effort not stated); Google Gemini 4 Argon launch benchmark table; methodology at deepmind.google/models/evals-methodology/gemini-4-argon; fallback not reported

    Metric: published score (%)

    Original source · Retrieved 2026-10-08

    Developer-reported aggregate. Same protocol as the catalog row: the lab table reproduces existing cells (gpt-6-astra 5.4, fable-5-1 6.7). Effort not stated.

  17. OSWorld 2.0 — OpenAI H2H (offline subset): 69.2%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Gemini 4 Argon (effort not stated); Google Gemini 4 Argon launch benchmark table; methodology at deepmind.google/models/evals-methodology/gemini-4-argon; fallback not reported

    Metric: published score (%)

    Original source · Retrieved 2026-10-08

    Developer-reported aggregate. Same protocol as the catalog row: the lab table reproduces existing cells (gpt-6-astra 72.6). Effort not stated. Offline subset, partial score, as labelled in the table.

Cite this model record

The Latent. “Gemini 4 Argon: score and benchmark evidence.” 2026-10-10. Release tli-2026-v1.0-2026-10-10-refresh-21.

Download release data and definitions (JSON) · Data usage terms