evaluation-task · Documentary review

PaperBench: Evaluating AI's Ability to Replicate AI Research (Work In Progress)

Agents are evaluated on their ability to replicate 20 ICML 2024 Spotlight and Oral papers from scratch. Given a research paper PDF, an addendum with clarifications, and a rubric defining evaluation criteria, the agent must reproduce the paper's key results by writing and executing code. > **Note:** This eval is a work in progress. See <https://github.com/UKGovernmentBEIS/inspect_evals/issues/334> for status.

Open interactive record →
OriginOpenAITopicsHuman influence & agency · Autonomy & agents · AI R&DStatusReviewed

Can support

How well an evaluated agent reproduces specified AI research outputs for the selected papers under the published PaperBench protocol.

Cannot support by itself

Unbounded original research ability, research taste, theory generation, lab management, reliable scientific autonomy, or acceleration of AI progress in deployment.

Decision use

Best used for

Comparing research-replication performance when paper set, rubric, judge, time, compute, tools, and scaffold are held fixed.

Not enough for

Claims that a system can autonomously conduct general AI research or replace an end-to-end research team.

Original sources