evaluation-task · Documentary review

StrongREJECT: Measuring LLM susceptibility to jailbreak attacks

A benchmark that evaluates the susceptibility of LLMs to various jailbreak attacks.

Open interactive record →
OriginAlexandra Souly, Qingyuan Lu, Dillon Bowen et al.TopicsSafeguardsStatusReviewed

Can support

How harmful and useful the target's responses are on the StrongREJECT prompt and attack distribution under the stated evaluator.

Cannot support by itself

Universal jailbreak robustness, real-world misuse success, adaptive agentic attacks, or downstream harm.

Decision use

Best used for

Evaluating jailbreak responses with a more discriminating rubric than simple refusal detection.

Not enough for

Certifying safety across threats, modalities, tools, model versions, or deployment contexts.

Original sources