evaluation-resource · Documentary review

Terminal-Bench

Evaluates agents completing real tasks in terminal environments.

Open interactive record →
OriginStanford and the Laude InstituteTopicsAutonomy & agents · CyberStatusReviewed

Can support

Ability to complete the benchmark's terminal tasks under a specified agent harness, environment, and budget.

Cannot support by itself

General computer-use autonomy, safe production administration, open-ended planning, or reliability across arbitrary terminal workloads.

Decision use

Best used for

Testing tool-using agents on reproducible terminal tasks and analysing failures at the trajectory level.

Not enough for

Claims of dependable autonomous software operation, cyber capability, or broad workplace automation.

Original sources