evaluation-package · Source-linked discovery

Do-Not-Answer

Evaluates LLM safety by presenting 939 prompts spanning five risk areas (information hazards, malicious uses, discrimination/toxicity, misinformation, human-chatbot harms) that responsible models should not directly answer. An LLM judge classifies each response into seven action categories (0-6), where category 6 indicates directly following the risky instruction. Metrics include refusal rate, safe rate (not category 6), and follow rate (category 6), broken down by risk area. Dataset was GPT-4-generated and human-filtered/annotated. An Inspect AI implementation of the original codebase publicly available at <https://github.com/Libr-AI/do-not-answer>.

Open interactive record →
OriginYuxia Wang, Haonan Li, Xudong Han et al.TopicsSafeguardsStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources