Gemini 3.8 Flash: score and benchmark evidence
Historical release · · Google
101.4 points · Rank 10 among 19 ranked models · Stability range 92.2–113.3 points.
Rank follows the point estimate; it does not establish statistically significant superiority. The scale is anchored at mean 100, SD 15 in the frozen calibration cohort, not human IQ or a percentage.
Release tli-2026-v4.1-2026-09-04. Methodology and limitations for this release.
This page preserves historical evidence. See the current model record.
Capability domains
Domain scores use calibration-cohort standard-deviation units. Unmeasured domains are shown as unavailable, not zero.
| Domain | Score (SD units) | Observed benchmarks |
|---|---|---|
| Knowledge & reasoning | 0.49 | 7 |
| Coding & software engineering | 0.35 | 7 |
| Agentic tool & computer use | 0.06 | 2 |
| Professional & real-world work | -0.47 | 12 |
| Multimodal & vision | -1.66 | 5 |
| Cybersecurity | Unavailable | 0 |
| Preference & communication | Unavailable | 0 |
Published benchmark evidence
33 results contribute to this score, from 33 collected results. Counts are not independent sample sizes.
This older release did not preserve per-model source records in its public snapshot. The archived scores remain available, but the current model record's evidence should not be treated as the source for these historical values.
Cite this model record
The Latent. “Gemini 3.8 Flash: score and benchmark evidence.” 2026-09-04. Release tli-2026-v4.1-2026-09-04.
Download release data and definitions (JSON) · Data usage terms
Related reporting
These dated stories provide reporting context; they may describe an earlier release.