evaluation-methodology · Documentary review
METR Time Horizons
Measures the length of software tasks frontier systems can complete with a given success probability.
OriginMETRTopicsAutonomy & agents · AI R&DStatusReviewed
Can support
A time-scaled summary of success on METR's selected task distribution under the stated elicitation, human-time estimates, and statistical model.
Cannot support by itself
Literal autonomous runtime, all economically useful work, general agent autonomy, safe long-term operation, reliability under distribution shift, or the fraction of jobs automatable.
Decision use
Best used for
Tracking a consistent slice of long-horizon task capability over model generations when task and elicitation methodology are sufficiently stable.
Not enough for
Universal autonomy thresholds, direct labor-automation forecasts, or deployment approval without reliability and safety evidence.