evaluation-package · Source-linked discovery

Do-Not-Answer Adversarial

An adversarial extension of the Do-Not-Answer benchmark that wraps each of the 939 harmful prompts (spanning five risk areas) in one of three jailbreak transforms -- prefix injection, refusal suppression, or their combination -- drawn from Wei et al. 2023. The model's response is scored against the original instruction using a seven-action behavioural taxonomy. The primary metric is follow_rate per risk area, with the delta versus the clean baseline indicating safeguard degradation under attack.

Open interactive record →
OriginYuxia Wang, Haonan Li, Xudong Han et al.TopicsSafeguardsStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources