evaluation-package · Documentary review

DeceptionBench

Measures belief-vs-behaviour deception in LLMs. For each of 180 scenarios the eval elicits the model's internal belief from a neutral prompt, then places the model under scenario pressure and captures its chain-of-thought (`<think>`) and final answer (`<output>`). Three verbatim GPT-4o judges check whether the reasoning tracks the belief (R~U_m) while the answer diverges from it (Y!~U_m). The headline metric is the Deception Tendency Rate (DTR) — the fraction of scenarios showing that deceptive pattern — reported overall and across five categories (sycophancy, strategic deception, honesty evasion, alignment faking, sandbagging).

Open interactive record →
OriginPKU-Alignment and the DeceptionBench paper authorsTopicsHuman influence & agency · Deception & misalignmentStatusReviewed

Can support

Deceptive response patterns under DeceptionBench's scenarios and operational definitions, conditional on the validity of its belief and reasoning proxies.

Cannot support by itself

A universal deception trait, private intent, spontaneous scheming, real-world deception prevalence, human susceptibility, or downstream harm.

Decision use

Best used for

Comparing protocol-defined deceptive behavior across scenarios and identifying sensitivity to incentives, prompts, and model conditions.

Not enough for

Claims that a model knowingly deceives in general or will deceive users in deployment.

Original sources