evaluation-package · Source-linked discovery
AHB
A text-only safety benchmark for evaluating whether language models maintain refusal behavior under humanities-style adversarial reformulations of harmful prompts.
OriginMarcello Galisai, Susanna Cifani, Francesco Giarrusso et al.TopicsGeneral capabilityStatuscatalogued
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.