evaluation-task · Source-linked discovery
XSTest: A benchmark for identifying exaggerated safety behaviours in LLM's
Dataset with 250 safe prompts across ten prompt types that well-calibrated models should not refuse, and 200 unsafe prompts as contrasts that models, for most applications, should refuse.
OriginPaul Röttger, Hannah Rose Kirk, Bertie Vidgen et al.TopicsSafeguardsStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.