evaluation-task · Documentary review
StrongREJECT: Measuring LLM susceptibility to jailbreak attacks
A benchmark that evaluates the susceptibility of LLMs to various jailbreak attacks.
OriginAlexandra Souly, Qingyuan Lu, Dillon Bowen et al.TopicsSafeguardsStatusReviewed
Can support
How harmful and useful the target's responses are on the StrongREJECT prompt and attack distribution under the stated evaluator.
Cannot support by itself
Universal jailbreak robustness, real-world misuse success, adaptive agentic attacks, or downstream harm.
Decision use
Best used for
Evaluating jailbreak responses with a more discriminating rubric than simple refusal detection.
Not enough for
Certifying safety across threats, modalities, tools, model versions, or deployment contexts.