Opus 5.5: score and benchmark evidence

Current published release · · Anthropic

144.2 points · Rank 1 among 25 ranked models · Stability range 133.3–155.2 points.

Rank follows the point estimate; it does not establish statistically significant superiority. The scale is anchored at mean 100, SD 15 in the frozen calibration cohort, not human IQ or a percentage.

One score summarizes demonstrated capability across published reasoning settings and seven equally weighted domains. Related benchmark variants share a family budget, and influence from one evaluator is capped.

Release tli-2026-v1.0-2026-09-22-gpt6pair-1. Methodology and limitations for this release.

Permanent link to these exact scores and sources

Capability domains

Domain scores use calibration-cohort standard-deviation units. Unmeasured domains are shown as unavailable, not zero.

DomainScore (SD units)Observed benchmarks
Knowledge & reasoning1.373
Coding & software engineering1.375
Agentic tool & computer use1.263
Professional & real-world work1.171
Multimodal & vision1.161
CybersecurityUnavailable0
Preference & communicationUnavailable0

Published benchmark evidence

13 results contribute to this score, from 40 collected results and 11 contributing benchmark families. Counts are not independent sample sizes.

Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.

The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.

Reporting sensitivity: ±19.8 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.

  1. SWE-bench Pro: 89.9%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 175 §8.2; retain the reported metric/version and reasoning setting.

  2. SWE-bench Multilingual — Opus 5.5 release: 93.9%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 175 §8.2; retain the reported metric/version and reasoning setting.

  3. SWE-bench Multimodal — Opus 5.5 release: 61.4%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 175 §8.2; retain the reported metric/version and reasoning setting.

  4. DeepSWE v1.1: 74.2%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 175 §8.3; retain the reported metric/version and reasoning setting.

  5. FrontierCode v1.1 Main — Anthropic H2H: 54.6%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (medium); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 176 §8.4; retain the reported metric/version and reasoning setting.

  6. FrontierCode v1.1 Extended: 65.3%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (medium); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 176 §8.4; retain the reported metric/version and reasoning setting.

  7. Terminal-Bench 4.0: 66.36%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (xhigh); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 178 §8.5; retain the reported metric/version and reasoning setting. xhigh=66.36 (rounded 66.4 in launch); fallback affects 10% of trials.

  8. Terminal-Bench Science 0.1: 58.7%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 178 §8.6; retain the reported metric/version and reasoning setting.

  9. FrontierSWE v2: 62.3%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 179 §8.7; retain the reported metric/version and reasoning setting.

  10. CursorBench 4.0: 57.8%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 179 §8.8; retain the reported metric/version and reasoning setting.

  11. MathArena — ArXivMath Aug 2026 — Anthropic no tools: 91.2%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; safeguard classifiers disabled; no fallback

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    System card p. 181–182 §8.9; retain the reported metric/version and reasoning setting.

  12. MathArena — ArXivMath Aug 2026 — Anthropic with tools: 96.9%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; safeguard classifiers disabled; no fallback

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    System card p. 181–183 §8.9; retain the reported metric/version and reasoning setting.

  13. ProgramBench — Opus 5.5 filtered test-pass rate: 91.2%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 183 §8.10.1; retain the reported metric/version and reasoning setting. Hidden-test pass rate on 166 filtered tasks, without six-hour timeout; not Vals raw task pass rate.

  14. Humanity's Last Exam (no tools): 64.4%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 174 table8.1.A; 184 §8.11.1; retain the reported metric/version and reasoning setting.

  15. Humanity's Last Exam w/ tools: 67.7%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 174 table8.1.A; 184 §8.11.1; retain the reported metric/version and reasoning setting.

  16. Chartography — Opus 5.5 no tools: 64.4%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 200 §8.13.1; retain the reported metric/version and reasoning setting.

  17. Chartography — Opus 5.5 with tools: 89%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 200 §8.13.1; retain the reported metric/version and reasoning setting.

  18. OSWorld 2.0 — Sep 10 tasks, Anthropic partial credit: 81.8%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 205–206 §8.13.3; retain the reported metric/version and reasoning setting. September 10 task assets plus server-side context management; explicitly incomparable to earlier releases.

  19. OSWorld 2.0 — Sep 10 tasks, Anthropic strict pass: 48.7%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 205–206 §8.13.3; retain the reported metric/version and reasoning setting. September 10 task assets plus server-side context management; explicitly incomparable to earlier releases.

  20. OfficeQA — Opus 5.5 release: 78.9%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 208 §8.14.1; retain the reported metric/version and reasoning setting.

  21. OfficeQA Pro †: 67.7%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 208 §8.14.1; retain the reported metric/version and reasoning setting.

  22. Toolathlon-Verified: 77.8%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; none; safeguard-stopped trials counted as failures

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    System card p. 210–211 §8.14.5; retain the reported metric/version and reasoning setting.

  23. AutomationBench — Anthropic H2H: 40%

    Contributes to the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; none; safeguard-stopped trials counted as failures

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    System card p. 174 table8.1.A; 211–212 §8.14.6; retain the reported metric/version and reasoning setting.

  24. HealthBench Professional — Opus 5.5 length-adjusted: 65.6%

    Not used in the score. Developer-reported. Review status: reviewed-developer.

    Configuration: Opus 5.5 (max); Anthropic system-card evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. System card p. 213–214 §8.15.2; retain the reported metric/version and reasoning setting. Opus 4.8 grader and length-adjusted metric; 77.1 raw is not substituted.

  25. AA-Briefcase v1.1 (Elo): 1821.85 native-points

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (max); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: Elo

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Exact model and effort in the evaluator chart payload. HLE is text-only/no tools. AutomationBench-AA is guardrail-adjusted objective credit, not raw objective completion or native whole-workflow pass rate.

  26. GDPval-AA v2.1 (Elo): 1846.17 native-points

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (max); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: Elo

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Exact model and effort in the evaluator chart payload. HLE is text-only/no tools. AutomationBench-AA is guardrail-adjusted objective credit, not raw objective completion or native whole-workflow pass rate.

  27. AutomationBench-AA: 69.5386812864%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (max); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Exact model and effort in the evaluator chart payload. HLE is text-only/no tools. AutomationBench-AA is guardrail-adjusted objective credit, not raw objective completion or native whole-workflow pass rate.

  28. Terminal-Bench 4.0 — AA harness: 59.595959596%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (xhigh); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Exact model and effort in the evaluator chart payload. HLE is text-only/no tools. AutomationBench-AA is guardrail-adjusted objective credit, not raw objective completion or native whole-workflow pass rate.

  29. SciCode: 66.8981481481%

    Contributes to the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (max); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Exact model and effort in the evaluator chart payload. HLE is text-only/no tools. AutomationBench-AA is guardrail-adjusted objective credit, not raw objective completion or native whole-workflow pass rate.

  30. HLE text-only — AA harness: 61.3531047266%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (max); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Exact model and effort in the evaluator chart payload. HLE is text-only/no tools. AutomationBench-AA is guardrail-adjusted objective credit, not raw objective completion or native whole-workflow pass rate.

  31. GDP.pdf — AA harness: 28.8%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (high); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Exact model and effort in the evaluator chart payload. HLE is text-only/no tools. AutomationBench-AA is guardrail-adjusted objective credit, not raw objective completion or native whole-workflow pass rate.

  32. CritPt: 31.7142857143%

    Contributes to the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (xhigh); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Exact model and effort in the evaluator chart payload. HLE is text-only/no tools. AutomationBench-AA is guardrail-adjusted objective credit, not raw objective completion or native whole-workflow pass rate.

  33. AA-Omniscience: 46.4166666667 native-points

    Contributes to the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (max); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: index points

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Exact model and effort in the evaluator chart payload. HLE is text-only/no tools. AutomationBench-AA is guardrail-adjusted objective credit, not raw objective completion or native whole-workflow pass rate.

  34. AA-LCR v1.1: 84.6666666667%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (xhigh); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Exact model and effort in the evaluator chart payload. HLE is text-only/no tools. AutomationBench-AA is guardrail-adjusted objective credit, not raw objective completion or native whole-workflow pass rate.

  35. MMMU-Pro: 87.6878612717%

    Contributes to the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (max); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Independent MMMU-Pro aggregate for this effort.

  36. Harvey LAB-AA: 91.243305833%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (xhigh); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Harvey is mean criterion-pass rate, not all-criteria task success.

  37. AA Intelligence Index v4.3.2 (score): 57.6223698103 native-points

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (max); Artificial Analysis independent evaluation, September 22 2026; production default fallback to older Claude models

    Metric: index points

    Original source · Retrieved 2026-09-22

    Scored as the shipped Opus 5.5 system. Default fallback can route cyber tasks to Opus 4.8 and biology/frontier-model-development tasks to Opus 5; this is not a pure-model claim. Reference composite only; not scored alongside its components.

  38. Box — Opus 5.5 Consumer products: 82%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (undisclosed); Box customer evaluation, September 22 2026; undisclosed

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Customer-run industry accuracy; sample counts and uncertainty undisclosed. Separate from existing Box overall benchmark; uncalibrated.

  39. Box — Opus 5.5 Financial services: 75%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (undisclosed); Box customer evaluation, September 22 2026; undisclosed

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Customer-run industry accuracy; sample counts and uncertainty undisclosed. Separate from existing Box overall benchmark; uncalibrated.

  40. Box — Opus 5.5 Technology: 74%

    Not used in the score. Source category: independent. Review status: reviewed-independent.

    Configuration: Opus 5.5 (undisclosed); Box customer evaluation, September 22 2026; undisclosed

    Metric: published benchmark score (%)

    Original source · Retrieved 2026-09-22

    Customer-run industry accuracy; sample counts and uncertainty undisclosed. Separate from existing Box overall benchmark; uncalibrated.

Cite this model record

The Latent. “Opus 5.5: score and benchmark evidence.” 2026-09-22. Release tli-2026-v1.0-2026-09-22-gpt6pair-1.

Download release data and definitions (JSON) · Data usage terms