evaluation-resource · Documentary review
HarmBench
Framework for automated red teaming and harmful-behaviour robustness.
OriginCenter for AI Safety and the HarmBench paper authorsTopicsSafeguardsStatusReviewed
Can support
Elicitability of HarmBench's harmful behaviors under the selected red-team methods and target configuration.
Cannot support by itself
Real-world harm, user intent, deployment prevalence, universal safeguard robustness, or agentic end-to-end misuse.
Decision use
Best used for
Comparing red-teaming methods, target-model vulnerability, and defense regressions on a standardized behavior set.
Not enough for
A comprehensive safety certification or estimate of incident probability.