evaluation-methodology · Documentary review

METR Time Horizons

Measures the length of software tasks frontier systems can complete with a given success probability.

Open interactive record →
OriginMETRTopicsAutonomy & agents · AI R&DStatusReviewed

Can support

A time-scaled summary of success on METR's selected task distribution under the stated elicitation, human-time estimates, and statistical model.

Cannot support by itself

Literal autonomous runtime, all economically useful work, general agent autonomy, safe long-term operation, reliability under distribution shift, or the fraction of jobs automatable.

Decision use

Best used for

Tracking a consistent slice of long-horizon task capability over model generations when task and elicitation methodology are sufficiently stable.

Not enough for

Universal autonomy thresholds, direct labor-automation forecasts, or deployment approval without reliability and safety evidence.

Original sources