evaluation-task · Source-linked discovery
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
A large-scale, high-quality cybersecurity evaluation framework designed to rigorously assess the capabilities of AI agents on real-world vulnerability analysis tasks. CyberGym includes 1,507 benchmark instances with historical vulnerabilities from 188 large software projects.
OriginZhun Wang, Tianneng Shi, Jingxuan He et al.TopicsCyberStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.