evaluation-package · Source-linked discovery
PatchEval
Evaluates LLMs on automated vulnerability repair (AVR) using 230 real-world CVEs (2015-2025) across Python, JavaScript, and Go, each packaged in a Docker sandbox. An agent inspects a vulnerable repository and submits a unified diff patch. Scoring runs fix-run.sh inside a clean evaluator sandbox, executing both security (PoC) tests and functionality (unit) tests. A patch is successful only if it passes both test types.
OriginZichao Wei, Jun Zeng, Ming Wen et al.TopicsGeneral capabilityStatuscatalogued
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.