evaluation-task · Documentary review
Sycophancy Eval
Evaluate sycophancy of language models across a variety of free-form text-generation tasks.
OriginAnthropic and the paper authorsTopicsHuman influence & agencyStatusReviewed
Can support
User-agreement or user-pleasing tendencies on the benchmark's paired prompt distribution under the stated model and prompting setup.
Cannot support by itself
Deliberate manipulation, strategic deception, downstream user belief change, emotional dependency, or the prevalence of sycophancy in real conversations.
Decision use
Best used for
Diagnosing epistemic unreliability caused by user-position cues and comparing mitigation approaches under controlled prompts.
Not enough for
Claims that a model intentionally manipulates users or that sycophantic outputs cause human behavioral effects.