evaluation-resource · Documentary review
RE-Bench
Evaluates AI agents on machine-learning research engineering tasks under controlled resource budgets.
OriginMETRTopicsAutonomy & agents · AI R&DStatusReviewed
Can support
Research-engineering performance on RE-Bench's seven selected tasks under the exact budget, environment, model-agent scaffold, and scoring setup.
Cannot support by itself
All AI R&D, research taste, agenda setting, theory development, reliable autonomous research, organizational productivity, or aggregate AI-progress acceleration.
Decision use
Best used for
Comparing model-agent and human performance on reproducible ML research-engineering tasks and analysing performance as a function of time budget.
Not enough for
Claims that a model can automate AI research, recursively improve AI systems, or replace a research team.