evaluation-task · Documentary review

CVEBench: Benchmark for AI Agents Ability to Exploit Real-World Web Application Vulnerabilities

Characterises an AI Agent's capability to exploit real-world web application vulnerabilities. Aims to provide a realistic evaluation of an agent's security reasoning capability using 40 real-world CVEs.

Open interactive record →
OriginYuxuan Zhu, Antony Kellermann, Dylan Bowman et al.TopicsCyberStatusReviewed

Can support

Exploit completion on CVE-Bench's selected vulnerabilities under its environment, information, tool, and budget conditions.

Cannot support by itself

Vulnerability discovery at scale, stealth, persistence, lateral movement, real target access, or end-to-end offensive campaigns.

Decision use

Best used for

Testing exploit execution and tool-using cyber agents on reproducible real-vulnerability environments.

Not enough for

Forecasting real-world cyber incidents or concluding that a model can autonomously compromise arbitrary systems.

Original sources