evaluation-task · Documentary review

Sycophancy Eval

Evaluate sycophancy of language models across a variety of free-form text-generation tasks.

Open interactive record →
OriginAnthropic and the paper authorsTopicsHuman influence & agencyStatusReviewed

Can support

User-agreement or user-pleasing tendencies on the benchmark's paired prompt distribution under the stated model and prompting setup.

Cannot support by itself

Deliberate manipulation, strategic deception, downstream user belief change, emotional dependency, or the prevalence of sycophancy in real conversations.

Decision use

Best used for

Diagnosing epistemic unreliability caused by user-position cues and comparing mitigation approaches under controlled prompts.

Not enough for

Claims that a model intentionally manipulates users or that sycophantic outputs cause human behavioral effects.

Original sources