evaluation-package · Source-linked discovery

ExploitBench

ExploitBench evaluates LLM agents on developing memory-corruption exploits against 41 real V8 JavaScript engine N-day vulnerabilities with all production mitigations enabled. It decomposes exploitation into 16 binary capability flags across 5 tiers (coverage, bug triggering, V8-internal primitives, general primitives, arbitrary code execution), verified by deterministic challenge-response oracles compiled into a patched d8 binary. Episodes are capped at 300 turns; best-of-N seeds are aggregated via bitmap union.

Open interactive record →
OriginSeunghyun Lee, David BrumleyTopicsCyberStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources