evaluation-task · Documentary review

APE: Attempt to Persuade Eval

Measures a model's willingness to attempt persuasion on harmful, controversial, and benign topics. The key metric is not persuasion effectiveness but whether the model attempts to persuade at all — particularly on harmful statements. Uses a multi-model setup: the evaluated model (persuader) converses with a simulated user (persuadee), and a third model (evaluator) scores each persuader turn for persuasion attempt. Based on the paper "It's the Thought that Counts" (arXiv:2506.02873).

Open interactive record →
OriginFAR AITopicsHuman influence & agencyStatusReviewed

Can support

Whether and how often the model attempts persuasion in APE's simulated dialogue distribution under the stated prompting and classification protocol.

Cannot support by itself

Persuasion effectiveness on humans, belief change, behavioral change, covert targeting, deployment at scale, durable agency loss, or electoral effects.

Decision use

Best used for

Comparing willingness to deploy persuasive strategies under a fixed set of simulated opportunities and model-system conditions.

Not enough for

Claims that a model is persuasive, manipulative, or capable of changing human behavior in real deployment.

Original sources