evaluation-task · Documentary review
PaperBench: Evaluating AI's Ability to Replicate AI Research (Work In Progress)
Agents are evaluated on their ability to replicate 20 ICML 2024 Spotlight and Oral papers from scratch. Given a research paper PDF, an addendum with clarifications, and a rubric defining evaluation criteria, the agent must reproduce the paper's key results by writing and executing code. > **Note:** This eval is a work in progress. See <https://github.com/UKGovernmentBEIS/inspect_evals/issues/334> for status.
Can support
How well an evaluated agent reproduces specified AI research outputs for the selected papers under the published PaperBench protocol.
Cannot support by itself
Unbounded original research ability, research taste, theory generation, lab management, reliable scientific autonomy, or acceleration of AI progress in deployment.
Decision use
Best used for
Comparing research-replication performance when paper set, rubric, judge, time, compute, tools, and scaffold are held fixed.
Not enough for
Claims that a system can autonomously conduct general AI research or replace an end-to-end research team.