evaluation-task · Documentary review
MASK: Disentangling Honesty from Accuracy in AI Systems
Evaluates honesty in large language models by testing whether they contradict their own beliefs when pressured to lie.
OriginCenter for AI SafetyTopicsHuman influence & agencyStatusReviewed
Can support
Whether model responses remain consistent with benchmark-elicited beliefs under MASK's controlled pressure scenarios and scoring rules.
Cannot support by itself
Private internal beliefs, privileged access to hidden cognition, persuasion effectiveness, human belief change, strategic deployment, or downstream social harm.
Decision use
Best used for
Studying protocol-specific honesty and motivated-misreporting behavior while separating it from ordinary factual accuracy more carefully than a single-answer test.
Not enough for
Claims that a model is generally truthful, intentionally deceptive in deployment, or effective at manipulating people.