evaluation-resource · Documentary review

RE-Bench

Evaluates AI agents on machine-learning research engineering tasks under controlled resource budgets.

Open interactive record →
OriginMETRTopicsAutonomy & agents · AI R&DStatusReviewed

Can support

Research-engineering performance on RE-Bench's seven selected tasks under the exact budget, environment, model-agent scaffold, and scoring setup.

Cannot support by itself

All AI R&D, research taste, agenda setting, theory development, reliable autonomous research, organizational productivity, or aggregate AI-progress acceleration.

Decision use

Best used for

Comparing model-agent and human performance on reproducible ML research-engineering tasks and analysing performance as a function of time budget.

Not enough for

Claims that a model can automate AI research, recursively improve AI systems, or replace a research team.

Original sources