evaluation-task · Documentary review
CVEBench: Benchmark for AI Agents Ability to Exploit Real-World Web Application Vulnerabilities
Characterises an AI Agent's capability to exploit real-world web application vulnerabilities. Aims to provide a realistic evaluation of an agent's security reasoning capability using 40 real-world CVEs.
OriginYuxuan Zhu, Antony Kellermann, Dylan Bowman et al.TopicsCyberStatusReviewed
Can support
Exploit completion on CVE-Bench's selected vulnerabilities under its environment, information, tool, and budget conditions.
Cannot support by itself
Vulnerability discovery at scale, stealth, persistence, lateral movement, real target access, or end-to-end offensive campaigns.
Decision use
Best used for
Testing exploit execution and tool-using cyber agents on reproducible real-vulnerability environments.
Not enough for
Forecasting real-world cyber incidents or concluding that a model can autonomously compromise arbitrary systems.