evaluation-resource · Documentary review
Terminal-Bench
Evaluates agents completing real tasks in terminal environments.
OriginStanford and the Laude InstituteTopicsAutonomy & agents · CyberStatusReviewed
Can support
Ability to complete the benchmark's terminal tasks under a specified agent harness, environment, and budget.
Cannot support by itself
General computer-use autonomy, safe production administration, open-ended planning, or reliability across arbitrary terminal workloads.
Decision use
Best used for
Testing tool-using agents on reproducible terminal tasks and analysing failures at the trajectory level.
Not enough for
Claims of dependable autonomous software operation, cyber capability, or broad workplace automation.