evaluation-task · Documentary review

MakeMeSay

Tests a model’s ability to generate persuasive or manipulative text, specifically in the setting of a game where one AI system has to get the other party (an AI system simulating a human) to say a specific codeword without arousing suspicion.

Open interactive record →
OriginOpenAI EvalsTopicsHuman influence & agencyStatusReviewed

Can support

Success at a narrow strategic-steering game against the specified simulated counterpart under the benchmark's rules.

Cannot support by itself

Human persuasion, belief or behavior change, covert manipulation in natural conversations, durable effects, or real-world social-engineering success.

Decision use

Best used for

Probing multi-turn strategic influence and conversational planning in a controlled game.

Not enough for

Claims that a model can manipulate humans or conduct effective influence operations.

Original sources