Sonnet 5.5: score and benchmark evidence
Current published release · · Anthropic
132.5 points · Rank 4 among 27 ranked models · Stability range 118.6–147 points.
Rank follows the point estimate; it does not establish statistically significant superiority. The scale is anchored at mean 100, SD 15 in the frozen calibration cohort, not human IQ or a percentage.
One score summarizes demonstrated capability across published reasoning settings and seven equally weighted domains. Related benchmark variants share a family budget, and influence from one evaluator is capped.
Release tli-2026-v1.0-2026-09-30-refresh-6. Methodology and limitations for this release.
Permanent link to these exact scores and sources
Capability domains
Domain scores use calibration-cohort standard-deviation units. Unmeasured domains are shown as unavailable, not zero.
| Domain | Score (SD units) | Observed benchmarks |
|---|---|---|
| Knowledge & reasoning | 0.94 | 1 |
| Coding & software engineering | 1.03 | 4 |
| Agentic tool & computer use | 0.99 | 3 |
| Professional & real-world work | 0.90 | 1 |
| Multimodal & vision | Unavailable | 0 |
| Cybersecurity | Unavailable | 0 |
| Preference & communication | Unavailable | 0 |
Published benchmark evidence
9 results contribute to this score, from 34 collected results and 7 contributing benchmark families. Counts are not independent sample sizes.
Results, source citations, review decisions and compatible reasoning modes are stored in a versioned relational database. Each publication uses a sealed, checksum-verified export; CSV files are optional exports, not editable scoring inputs. Previous revisions remain available for audit. Collected means a numeric result exists in the database. Used means its source, exact model checkpoint and benchmark protocol have been reviewed and its benchmark has a frozen calibration supported by multiple providers. For each benchmark we select the highest published aggregate across comparable reasoning modes, not necessarily maximum effort, and count it once. More tested settings provide more chances to record a high score; this selection advantage is not corrected by the conditional interval. Developer-reported and sole/unspecified settings remain eligible and disclosed. Unresolved does not mean false or useless: conflicting metrics, unverified citations and unmatched versions stay collected until reconciled.
The score stability range is a 95% bootstrap score interval conditional on frozen calibration and comparable evidence availability. It resamples benchmark families and evaluators and adds a frozen missing-evidence correction. It excludes selection across reasoning settings, selective publication, calibration changes and new domains. It is not a 95% guarantee about true capability or where a future score will land. Sparse or conflicting evidence can widen the range.
Reporting sensitivity: ±23.6 points. Sensitivity to selectively published results. This stress range is not a confidence interval or a probability statement.
SWE-bench Pro: 81.3%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row.
SWE-bench Multilingual — Opus 5.5 release: 90.3%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row.
SWE-bench Multimodal — Opus 5.5 release: 54.3%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row.
DeepSWE v1.1: 71%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row.
FrontierCode v1.1 Main — Anthropic H2H: 52.1%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (xhigh); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. Cognition ran v1.1 in Claude Code; official mergeability score. xhigh exceeds max. Operator data archive did not yet contain this model.
FrontierCode v1.1 Extended: 64.4%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (xhigh); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. Cognition ran v1.1 in Claude Code; official mergeability score. xhigh exceeds max. Operator data archive did not yet contain this model.
Terminal-Bench 4.0: 70.6%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. Claude Code --bare, no internet; fallback affected 1.5% of trials (1.2% of requests). Mean of five trials per task.
Terminal-Bench Science 0.1: 59.9%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; safeguards enabled; no fallback fired
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row.
FrontierSWE v2: 61.9%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. Measured by Proximal, 34 tasks, mean of five trials per task, as reported by Anthropic.
CursorBench 4.0: 55.5%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. Cursor measured results in its production harness and supplied them to Anthropic; developer-hosted report, not a separately retrieved operator result.
MathArena — ArXivMath Aug 2026 — Anthropic no tools: 86.8%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; safeguard classifiers disabled; no fallback
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. August 57-problem set, four attempts per problem; not MathArena agentic leaderboard protocol.
MathArena — ArXivMath Aug 2026 — Anthropic with tools: 95.2%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; safeguard classifiers disabled; no fallback
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. August 57-problem set, four attempts per problem; not MathArena agentic leaderboard protocol.
ProgramBench — Opus 5.5 filtered test-pass rate: 79.7%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. 166 filtered tasks, hidden-test pass rate, no six-hour timeout; not Vals raw task pass rate.
Humanity's Last Exam (no tools): 56.9%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row.
Humanity's Last Exam w/ tools: 64.5%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row.
Chartography — Opus 5.5 no tools: 61.6%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. 100 tasks, Gemini 3.5 Flash grader, five runs; same Opus 5.5 release protocol.
Chartography — Opus 5.5 with tools: 90.2%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. 100 tasks, Gemini 3.5 Flash grader, five runs; same Opus 5.5 release protocol.
OSWorld 2.0 — Sep 10 tasks, Anthropic partial credit: 80.1%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. v2.1 September 10 assets, 108 tasks, server-side compaction; card explicitly says unchanged from Opus 5.5 configuration. Partial credit and strict pass stay separate.
OSWorld 2.0 — Sep 10 tasks, Anthropic strict pass: 43.5%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. v2.1 September 10 assets, 108 tasks, server-side compaction; card explicitly says unchanged from Opus 5.5 configuration. Partial credit and strict pass stay separate.
OfficeQA — Opus 5.5 release: 76.9%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row.
OfficeQA Pro †: 65.6%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row.
Harvey LAB-AA: 93.1%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. AA harness, 120 held-out tasks, mean criterion pass. Do not substitute all-criteria task success.
GDPval-AA v2.1 (Elo): 1844 native-points
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: Elo
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. Artificial Analysis independent evaluation as reported by Anthropic. Pre-release serving of launch checkpoint had a structured-output bug, since fixed; do not infer a post-fix score.
AA-Briefcase v1.1 (Elo): 1811 native-points
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: Elo
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. Artificial Analysis independent evaluation as reported by Anthropic. Pre-release serving of launch checkpoint had a structured-output bug, since fixed; do not infer a post-fix score.
Toolathlon-Verified: 77.8%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; none; safeguard-stopped trials counted as failures
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. Pass@1 across 3 trials × 108 tasks; not Pass@3 or Pass³.
AutomationBench — Anthropic H2H: 44.7%
Contributes to the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. Zapier private evaluation set v1.0.6, whole-workflow success with default fallback; separate from public set and AA partial credit. Do not update Opus from the selective refusal-only rerun.
HealthBench Professional — Opus 5.5 length-adjusted: 69.2%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Same evaluation protocol as the corresponding catalog row. Opus 4.8 grader, length-adjusted metric; 77.1 raw is separate. Same Opus 5.5 baseline and adjustment formula.
BenchCAD — Anthropic 1000-file Vision2Code no tools: 0.747 native-points
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: mean voxel IoU
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Mean voxel IoU over five runs; raw 0–1 units, not percent task success. Kept separate from the older modified-settings comparison.
BenchCAD — Anthropic 1000-file Vision2Code with tools: 0.963 native-points
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: mean voxel IoU
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Mean voxel IoU over five runs; raw 0–1 units, not percent task success. Kept separate from the older modified-settings comparison.
Harvey LAB — AA held-out all-pass: 11.7%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (high); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. 120 held-out tasks, AA harness, all-criteria success; separate from mean criterion pass and older Harvey harnesses.
HealthBench — Anthropic Opus 4.8 grader length-adjusted: 65.4%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Exact labels in figure; distinct from raw and OpenAI-graded results.
HealthBench — Anthropic Opus 4.8 grader raw: 69.4%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Raw score before response-length adjustment; retained separately and uncalibrated.
HealthBench Professional — Anthropic Opus 4.8 grader raw: 77.1%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. Raw score before response-length adjustment; retained separately and uncalibrated.
PhysicianBench — Anthropic Opus 5 grader: 63.2%
Not used in the score. Developer-reported. Review status: reviewed-developer.
Configuration: Sonnet 5.5 (max); Anthropic September 28 system card; production cyber fallback to Sonnet 5; biology requests may stop
Metric: published score (%)
Original source · Retrieved 2026-09-28
Exact Sonnet 5.5 published aggregate. Shipped-system results can include disclosed Sonnet 5 fallback; no predecessor score is copied. 100 EHR tasks, pass@1, same agent harness and Opus 5 grading; exact figure labels.
Cite this model record
The Latent. “Sonnet 5.5: score and benchmark evidence.” 2026-09-30. Release tli-2026-v1.0-2026-09-30-refresh-6.
Download release data and definitions (JSON) · Data usage terms