The Latent
EN▾
EnglishEspañol中文PortuguêsFrançaisالعربية日本語한국어
Sign Up
NEWSBENCHMARKSDATALEARNNEWSLETTERPARTNER WITH US →
Data/Capability

Capability

METR AI Task-Completion Time Horizons

METR's estimated human-expert task duration at which frontier agents succeed 50% of the time, in hours. The metric provides task-suite evidence on agents' ability to complete longer work, relevant to deployment and oversight; TH1 and TH1.1 remain separate, and results above 16 hours are unreliable.

METR AI Task-Completion Time Horizons

  • TH1.1 50% horizon estimate
  • TH1.1 95% CI lower
  • TH1.1 95% CI upper (16h cap)
  • TH1 50% horizon estimate
  • TH1 95% CI lower
  • TH1 95% CI upper (16h cap)
The latest raw-data-backed TH1.1 frontier estimate is 16 hours hours for the model released on 2026-02-05; values and confidence bounds are displayed only through the 16-hour reliability ceiling.1001010.1Feb '25Apr '25Jun '25Aug '25Oct '25Dec '25Feb '26{"f":[800,420,56,16],"s":[["TH1.1 50% horizon estimate","#2563EB"],["TH1.1 95% CI lower","#60A5FA"],["TH1.1 95% CI upper (16h cap)","#93C5FD"],["TH1 50% horizon estimate","#475569"],["TH1 95% CI lower","#64748B"],["TH1 95% CI upper (16h cap)","#94A3B8"]],"p":[["2025-02-24T00:00:00.000Z","Feb '25",56,[[0,"1.01 hours",258.33,null],[1,"0.56 hours",289.56,null],[2,"1.79 hours",228.03,null],[3,"0.92 hours",263,null],[4,"0.48 hours",297.21,null],[5,"1.49 hours",237.78,null]],null],["2025-04-16T00:00:00.000Z","Apr '25",163.31,[[0,"2 hours",222.26,null],[1,"1.22 hours",248.35,null],[2,"3.19 hours",197.49,null],[3,"1.52 hours",236.46,null],[4,"0.78 hours",272.02,null],[5,"2.63 hours",207.63,null]],null],["2025-07-09T00:00:00.000Z","Jul '25",340.05,[[3,"1.75 hours",229.25,null],[4,"0.79 hours",270.92,null],[5,"3.47 hours",193.08,null]],null],["2025-08-05T00:00:00.000Z","Aug '25",396.86,[[3,"1.82 hours",227.08,null],[4,"0.93 hours",262.34,null],[5,"3.35 hours",195.01,null]],null],["2025-08-07T00:00:00.000Z","Aug '25",401.06,[[0,"3.38 hours",194.44,null],[1,"1.9 hours",224.75,null],[2,"6.78 hours",157.82,null],[3,"2.18 hours",217.59,null],[4,"1.13 hours",252.3,null],[5,"4.13 hours",183.93,null]],null],["2025-11-18T00:00:00.000Z","Nov '25",617.78,[[0,"3.74 hours",189.18,null],[1,"2.28 hours",215.21,null],[2,"6.46 hours",160.38,null]],null],["2025-11-19T00:00:00.000Z","Nov '25",619.88,[[3,"2.7 hours",206.25,null],[4,"1.34 hours",243.38,null],[5,"5.61 hours",167.81,null]],null],["2025-11-24T00:00:00.000Z","Nov '25",630.4,[[0,"4.88 hours",175.1,null],[1,"2.68 hours",206.81,null],[2,"10.64 hours",134.05,null],[3,"4.11 hours",184.15,null],[4,"1.74 hours",229.55,null],[5,"13.69 hours",120.77,null]],null],["2025-12-11T00:00:00.000Z","Dec '25",666.17,[[0,"5.87 hours",165.4,null],[1,"3.19 hours",197.56,null],[2,"14.37 hours",118.22,null]],null],["2026-02-05T00:00:00.000Z","Feb '26",784,[[0,"11.98 hours",127.81,null],[1,"5.32 hours",170.57,null],[2,"16 hours",112.57,null]],null]]}TH1.1 50% horizon estimate: 12 on Feb '26. TH1.1 95% CI lower: 5.3 on Feb '26. TH1.1 95% CI upper (16h cap): 16 on Feb '26. TH1 50% horizon estimate: 4.1 on Nov '25. TH1 95% CI lower: 1.7 on Nov '25. TH1 95% CI upper (16h cap): 13.7 on Nov '25
SOURCE: METR eval-analysis-public
RANGEALLYTD12M3M1M
Key takeaway

In METR's latest published GitHub raw data, Claude Opus 4.6 has a TH1.1 50% time-horizon estimate of 11.98 hours, indicating success extends to longer tasks within this suite. Its bootstrapped 95% interval is 5.32–65.83 hours; the chart clips the upper bound at 16 hours because METR says measurements above that threshold are unreliable.

The latest raw-data-backed TH1.1 frontier estimate is 16 hours hours for the model released on 2026-02-05; values and confidence bounds are displayed only through the 16-hour reliability ceiling.

Pro API coming soon

Methodology

The source observations are METR's public TH1 and TH1.1 runs.jsonl files at Git commit 52cb829c7a2efb2d659285c4b1768d191d97f8d2. The chart was reproduced from those raw runs using the analysis code and parameters in the same commit.

For each model, METR fits logistic regression to success as a function of log2 human_minutes, using the diversity-adjusted invsqrt_task_weight, regularization 0.00001 and the 50% success crossing as the time horizon. Minutes are divided by 60 for this chart.

The 95% confidence limits are the 2.5th and 97.5th percentiles from 1,000 deterministic hierarchical bootstrap samples over task family, task and run, matching METR's public pipeline.

Only frontier models are retained. Following METR's code, a model is frontier at release when its 50% horizon is at least the highest estimate among models released on or before that date. Human rows and non-frontier model rows are excluded.

Timestamps are model release dates from METR's release_dates.yaml, not evaluation or publication dates. There are no synthetic dates, interpolated observations or trend-line estimates in the CSV. Lines may visually connect observations within one benchmark version only.

TH1 and TH1.1 are not merged into one continuous series. TH1.1 changed both the task suite and evaluation infrastructure, so all estimates and confidence bounds are separately labeled. METR reported that the new estimates generally fall within TH1 confidence intervals but should not be treated as directly interchangeable.

METR warns that measurements above 16 hours are unreliable because the current suite does not adequately constrain very long-task performance. All point estimates included here are at or below 16 hours. To keep every plotted value within the requested ceiling, Claude Opus 4.6's TH1.1 upper confidence limit is visibly capped at 16 hours; its uncapped 65.829173-hour bound remains in the CSV note and key takeaway.

The repository's latest raw TH1.1 observation is for models released on February 5, 2026. METR's live page lists later 2026 updates, but corresponding later raw runs were not present in the official GitHub repository when checked through August 27, 2026, so they are not added or inferred.

Frequently asked questions

What does a 50% task-completion time horizon mean?

It is the human-expert duration of tasks at which METR's fitted curve predicts the AI agent will succeed half the time. It measures task difficulty in human time, not how long the AI itself runs and not the duration of every task the model can complete.

Why does the time horizon matter to the AI industry?

It provides a standardized view of whether frontier agents can reliably complete tasks that take human experts progressively longer, which is relevant to assessing useful autonomy, deployment scope and oversight needs. It is evidence for METR's task distribution, not a forecast of job automation or performance in every workplace.

How should a rising horizon be interpreted?

Within one methodology version, a higher estimate means the fitted 50% success point occurs on tasks with longer human-expert completion times. The increase can reflect better model and agent performance on this suite, but it does not specify which skills improved or guarantee reliability on a particular long task.

Why are TH1 and TH1.1 separate lines?

TH1.1 expanded and revised the task suite and moved evaluations from Vivaria to Inspect. METR retrospectively re-estimated a subset of models, but the methodology versions measure somewhat different task distributions and must not be joined into one interpolated line.

Why does the chart stop at 16 hours?

METR says measurements above 16 hours are unreliable with its current task suite. The chart therefore excludes point estimates above that threshold and clips the single upper confidence bound that exceeds it, while disclosing the uncapped bound.

Does this include METR results added after February 2026?

No. The official GitHub raw files at the latest available repository revision end with models released on February 5, 2026. Later models shown on METR's live page are omitted because matching GitHub raw runs were unavailable as of August 27, 2026.

How were the confidence intervals calculated?

They are 95% percentile intervals from 1,000 hierarchical bootstrap samples over task families, tasks and individual runs, using METR's public analysis implementation.

What data and method underlie the estimates?

The chart reproduces METR's TH1 and TH1.1 raw runs and public analysis code at a pinned Git commit. For each model, logistic regression relates success to log2 human task time; the 50% crossing is converted from minutes to hours, and only models on METR's release-date frontier are retained.

Related charts

  • AI Benchmark Saturation Curves
  • Largest Documented Training Compute by Release Year: China vs. U.S.
  • Frontier Model Training Compute

The Latent

AI industry news. A sister publication to The Block.

Editorial

  • Standards
  • Corrections
  • Commercial policy
  • Contact

Company

  • LEARN
  • Data
  • Benchmarks
  • Score
  • Methodology
  • About
  • Team
  • Privacy Policy
  • Terms of Service
  • Security
  • The Block
  • Add The Latent as a preferred source