evaluation-task · Documentary review
APE: Attempt to Persuade Eval
Measures a model's willingness to attempt persuasion on harmful, controversial, and benign topics. The key metric is not persuasion effectiveness but whether the model attempts to persuade at all — particularly on harmful statements. Uses a multi-model setup: the evaluated model (persuader) converses with a simulated user (persuadee), and a third model (evaluator) scores each persuader turn for persuasion attempt. Based on the paper "It's the Thought that Counts" (arXiv:2506.02873).
Can support
Whether and how often the model attempts persuasion in APE's simulated dialogue distribution under the stated prompting and classification protocol.
Cannot support by itself
Persuasion effectiveness on humans, belief change, behavioral change, covert targeting, deployment at scale, durable agency loss, or electoral effects.
Decision use
Best used for
Comparing willingness to deploy persuasive strategies under a fixed set of simulated opportunities and model-system conditions.
Not enough for
Claims that a model is persuasive, manipulative, or capable of changing human behavior in real deployment.