evaluation-family · Documentary review

Agentharm

Assesses whether AI agents might engage in harmful activities by testing their responses to malicious prompts in areas like cybercrime, harassment, and fraud, aiming to ensure safe behavior.

Open interactive record →
OriginGray Swan AI and the UK AI Security InstituteTopicsAutonomy & agents · SafeguardsStatusReviewed

Can support

Whether an evaluated agent can be elicited to execute AgentHarm's selected harmful tasks under the published tools, environment, and jailbreak setup.

Cannot support by itself

Real-world misuse prevalence, operator skill, deployment access, victim impact, or the probability that a deployed agent will cause harm.

Decision use

Best used for

Stress-testing harmful agent behavior and the interaction between refusal safeguards and retained task capability.

Not enough for

Estimating incident rates or certifying that an agent is safe in operational environments.

Original sources