evaluation-package · Source-linked discovery
Alignment Faking
Measures whether LLMs strategically comply with harmful requests when believing outputs will be used for RLHF training (free-tier/monitored) versus when unmonitored (paid-tier). Uses 100 harmful prompts from LLM-LAT/harmful-dataset. Each prompt is presented twice with free-tier and paid-tier tags inside a system prompt describing a conflicting training objective. Primary metric is the compliance gap (free minus paid compliance rate via two-proportion z-test); a secondary scorer detects verbalized alignment-faking reasoning in scratchpad outputs.
OriginRyan Greenblatt, Carson Denison, Benjamin Wright et al.TopicsDeception & misalignmentStatuscatalogued
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.