evaluation-task · Documentary review

MASK: Disentangling Honesty from Accuracy in AI Systems

Evaluates honesty in large language models by testing whether they contradict their own beliefs when pressured to lie.

Open interactive record →
OriginCenter for AI SafetyTopicsHuman influence & agencyStatusReviewed

Can support

Whether model responses remain consistent with benchmark-elicited beliefs under MASK's controlled pressure scenarios and scoring rules.

Cannot support by itself

Private internal beliefs, privileged access to hidden cognition, persuasion effectiveness, human belief change, strategic deployment, or downstream social harm.

Decision use

Best used for

Studying protocol-specific honesty and motivated-misreporting behavior while separating it from ordinary factual accuracy more carefully than a single-answer test.

Not enough for

Claims that a model is generally truthful, intentionally deceptive in deployment, or effective at manipulating people.

Original sources