evaluation-package · Source-linked discovery

AHB

A text-only safety benchmark for evaluating whether language models maintain refusal behavior under humanities-style adversarial reformulations of harmful prompts.

Open interactive record →
OriginMarcello Galisai, Susanna Cifani, Francesco Giarrusso et al.TopicsGeneral capabilityStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources