evaluation-package · Documentary review
DeceptionBench
Measures belief-vs-behaviour deception in LLMs. For each of 180 scenarios the eval elicits the model's internal belief from a neutral prompt, then places the model under scenario pressure and captures its chain-of-thought (`<think>`) and final answer (`<output>`). Three verbatim GPT-4o judges check whether the reasoning tracks the belief (R~U_m) while the answer diverges from it (Y!~U_m). The headline metric is the Deception Tendency Rate (DTR) — the fraction of scenarios showing that deceptive pattern — reported overall and across five categories (sycophancy, strategic deception, honesty evasion, alignment faking, sandbagging).
Can support
Deceptive response patterns under DeceptionBench's scenarios and operational definitions, conditional on the validity of its belief and reasoning proxies.
Cannot support by itself
A universal deception trait, private intent, spontaneous scheming, real-world deception prevalence, human susceptibility, or downstream harm.
Decision use
Best used for
Comparing protocol-defined deceptive behavior across scenarios and identifying sensitivity to incentives, prompts, and model conditions.
Not enough for
Claims that a model knowingly deceives in general or will deceive users in deployment.