evaluation-resource · Documentary review

HarmBench

Framework for automated red teaming and harmful-behaviour robustness.

Open interactive record →
OriginCenter for AI Safety and the HarmBench paper authorsTopicsSafeguardsStatusReviewed

Can support

Elicitability of HarmBench's harmful behaviors under the selected red-team methods and target configuration.

Cannot support by itself

Real-world harm, user intent, deployment prevalence, universal safeguard robustness, or agentic end-to-end misuse.

Decision use

Best used for

Comparing red-teaming methods, target-model vulnerability, and defense regressions on a standardized behavior set.

Not enough for

A comprehensive safety certification or estimate of incident probability.

Original sources