evaluation-task · Documentary review

Agentic Misalignment: How LLMs could be insider threats

Eliciting unethical behaviour (most famously blackmail) in response to a fictional company-assistant scenario where the model is faced with replacement.

Open interactive record →
OriginAnthropicTopicsDeception & misalignment · Autonomy & agentsStatusReviewed

Can support

Harmful action selection in the benchmark's synthetic high-conflict scenarios under a specified model, prompt, tool, and opportunity configuration.

Cannot support by itself

The base rate of misalignment in normal deployment, spontaneous long-term goals, real organizational incidents, or the probability of autonomous insider threats.

Decision use

Best used for

Stress-testing model systems under explicit conflict and opportunity and comparing scenario or safeguard interventions.

Not enough for

Estimating real-world incident probability or concluding that a model has persistent malicious goals.

Original sources