MiMo-V2.6-Pro: score and benchmark evidence
Historical release · · Xiaomi
125.8 points · Rank 5 among 25 ranked models · Stability range 112–138.3 points.
Rank follows the point estimate; it does not establish statistically significant superiority. The scale is anchored at mean 100, SD 15 in the frozen calibration cohort, not human IQ or a percentage.
One score summarizes demonstrated capability across published reasoning settings and seven equally weighted domains. Related benchmark variants share a family budget, and influence from one evaluator is capped.
Release tli-2026-v1.0-2026-09-22-gpt6pair-1. Methodology and limitations for this release.
This page preserves historical evidence. See the current model record.
Capability domains
Domain scores use calibration-cohort standard-deviation units. Unmeasured domains are shown as unavailable, not zero.
| Domain | Score (SD units) | Observed benchmarks |
|---|---|---|
| Knowledge & reasoning | 0.61 | 3 |
| Coding & software engineering | 0.87 | 4 |
| Agentic tool & computer use | 0.53 | 3 |
| Professional & real-world work | 0.69 | 1 |
| Multimodal & vision | Unavailable | 0 |
| Cybersecurity | 0.98 | 1 |
| Preference & communication | Unavailable | 0 |
Published benchmark evidence
12 results contribute to this score, from 29 collected results and 11 contributing benchmark families. Counts are not independent sample sizes.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
Reporting sensitivity: ±19.8 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
DeepSWE v1.1: 71.9%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Released benchmark table 71.9; training-progress endpoint 72.57 is not the released checkpoint result and is not substituted.
ProgramBench — MiMo V2.6 release: 26.5%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
MiMo Code Bench: 63.2%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
AutomationBench v1.0.6: 53.1%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
Toolathlon-Verified: 76.9%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
Agents' Last Exam: 31.6%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
Terminal-Bench 4.0: 34.9%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
Terminal-Bench 2.1: 89.9%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
OSWorld-Verified: 82%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
JobBench: 62%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
CyberGym: 94%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
MiMo Cyber Bench: 80.2%
Not used in the score. Developer-reported. Review status: unresolved.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Unresolved: model card reports 80.2, while the official release bench.js reports 81.7. Neither is admitted to the score.
ExploitGym: 17.8%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
ExploitBench — MiMo V2.6 release: 47.9%
Not used in the score. Developer-reported. Review status: unresolved.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Release table reports 47.9; supplied source review identifies a conflicting product chart value of 43.8. Protocol/version equivalence and discrepancy remain unresolved; not scored.
SEC-Bench Pro — MiMo V2.6 release: 66.3%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
MiMo Visual Coding: 72.3%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: MiMo-V2.6-Pro; Xiaomi release evaluation; harness undisclosed; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Reviewed published release aggregate; no effort or harness equivalence inferred for versioned/ambiguous protocols.
AA-Briefcase v1.1 (Elo): 1521.82 native-points
Not used in the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; Artificial Analysis independent evaluation, September 22 2026; effort, retries and fallback undisclosed
Metric: Elo
Original source · Retrieved 2026-09-22
Directly reviewed exact-model chart payload from the independent evaluator; not a reseller attribution.
GDPval-AA v2.1 (Elo): 1673.19 native-points
Not used in the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; Artificial Analysis independent evaluation, September 22 2026; effort, retries and fallback undisclosed
Metric: Elo
Original source · Retrieved 2026-09-22
Directly reviewed exact-model chart payload from the independent evaluator; not a reseller attribution. Native 1673.19 Elo corroborates rounded 1673 from Xiaomi; normalized 58.6595% is not an additional observation.
AutomationBench-AA: 58.6234094342%
Not used in the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; Artificial Analysis independent evaluation, September 22 2026; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Directly reviewed exact-model chart payload from the independent evaluator; not a reseller attribution.
Terminal-Bench 4.0 — AA harness: 34.8484848485%
Not used in the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; Artificial Analysis independent evaluation, September 22 2026; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Directly reviewed exact-model chart payload from the independent evaluator; not a reseller attribution.
SciCode: 60.8796296296%
Contributes to the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; Artificial Analysis independent evaluation, September 22 2026; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Directly reviewed exact-model chart payload from the independent evaluator; not a reseller attribution.
Humanity's Last Exam (no tools): 49.3512511585%
Contributes to the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; Artificial Analysis independent evaluation, September 22 2026; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Directly reviewed exact-model chart payload from the independent evaluator; not a reseller attribution.
GDP.pdf — AA harness: 19.2%
Not used in the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; Artificial Analysis independent evaluation, September 22 2026; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Directly reviewed exact-model chart payload from the independent evaluator; not a reseller attribution.
CritPt: 26.5714285714%
Contributes to the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; Artificial Analysis independent evaluation, September 22 2026; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Directly reviewed exact-model chart payload from the independent evaluator; not a reseller attribution.
AA-Omniscience: 8.3833333333 native-points
Contributes to the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; Artificial Analysis independent evaluation, September 22 2026; effort, retries and fallback undisclosed
Metric: index points
Original source · Retrieved 2026-09-22
Directly reviewed exact-model chart payload from the independent evaluator; not a reseller attribution. Index 8.3833333333; accuracy 34.85% and non-hallucination 59.3757994372% are different metrics, not the index and not additional scored benchmarks.
AA-LCR v1.1: 86.3333333333%
Not used in the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; Artificial Analysis independent evaluation, September 22 2026; effort, retries and fallback undisclosed
Metric: published benchmark score (%)
Original source · Retrieved 2026-09-22
Directly reviewed exact-model chart payload from the independent evaluator; not a reseller attribution.
AA Intelligence Index v4.3.2 (score): 46.324206531 native-points
Not used in the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; Artificial Analysis independent evaluation, September 22 2026; effort, retries and fallback undisclosed
Metric: index points
Original source · Retrieved 2026-09-22
Composite retained for reference only; not substituted into v4.1.1 and not scored alongside its components.
OpenLM Arena+ — Overall (Elo): 1507 native-points
Not used in the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; OpenLM Arena+ LLM-as-a-judge; overall track; effort, retries and fallback undisclosed
Metric: Elo
Original source · Retrieved 2026-09-22
Direct operator listing. LLM-as-a-judge, not human-vote LMArena; shares a suite with the coding track.
OpenLM Arena+ — Coding (Elo): 1560 native-points
Not used in the score. Source category: independent. Review status: reviewed-independent.
Configuration: MiMo-V2.6-Pro; OpenLM Arena+ LLM-as-a-judge; coding track; effort, retries and fallback undisclosed
Metric: Elo
Original source · Retrieved 2026-09-22
Direct operator coding-track listing; LLM-as-a-judge. Shares a suite with the overall track; no calibrated score contribution.
Cite this model record
The Latent. “MiMo-V2.6-Pro: score and benchmark evidence.” 2026-09-22. Release tli-2026-v1.0-2026-09-22-gpt6pair-1.
Download release data and definitions (JSON) · Data usage terms