evaluation-package · Source-linked discovery
Do-Not-Answer Adversarial
An adversarial extension of the Do-Not-Answer benchmark that wraps each of the 939 harmful prompts (spanning five risk areas) in one of three jailbreak transforms -- prefix injection, refusal suppression, or their combination -- drawn from Wei et al. 2023. The model's response is scored against the original instruction using a seven-action behavioural taxonomy. The primary metric is follow_rate per risk area, with the delta versus the clean baseline indicating safeguard degradation under attack.
OriginYuxia Wang, Haonan Li, Xudong Han et al.TopicsSafeguardsStatuscatalogued
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.