evaluation-task · Documentary review
Agentic Misalignment: How LLMs could be insider threats
Eliciting unethical behaviour (most famously blackmail) in response to a fictional company-assistant scenario where the model is faced with replacement.
OriginAnthropicTopicsDeception & misalignment · Autonomy & agentsStatusReviewed
Can support
Harmful action selection in the benchmark's synthetic high-conflict scenarios under a specified model, prompt, tool, and opportunity configuration.
Cannot support by itself
The base rate of misalignment in normal deployment, spontaneous long-term goals, real organizational incidents, or the probability of autonomous insider threats.
Decision use
Best used for
Stress-testing model systems under explicit conflict and opportunity and comparing scenario or safeguard interventions.
Not enough for
Estimating real-world incident probability or concluding that a model has persistent malicious goals.