evaluation-package · Source-linked discovery

Alignment Faking

Measures whether LLMs strategically comply with harmful requests when believing outputs will be used for RLHF training (free-tier/monitored) versus when unmonitored (paid-tier). Uses 100 harmful prompts from LLM-LAT/harmful-dataset. Each prompt is presented twice with free-tier and paid-tier tags inside a system prompt describing a conflicting training objective. Primary metric is the compliance gap (free minus paid compliance rate via two-proportion z-test); a secondary scorer detects verbalized alignment-faking reasoning in scratchpad outputs.

Open interactive record →
OriginRyan Greenblatt, Carson Denison, Benjamin Wright et al.TopicsDeception & misalignmentStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources