evaluation-family · Documentary review
Agentharm
Assesses whether AI agents might engage in harmful activities by testing their responses to malicious prompts in areas like cybercrime, harassment, and fraud, aiming to ensure safe behavior.
OriginGray Swan AI and the UK AI Security InstituteTopicsAutonomy & agents · SafeguardsStatusReviewed
Can support
Whether an evaluated agent can be elicited to execute AgentHarm's selected harmful tasks under the published tools, environment, and jailbreak setup.
Cannot support by itself
Real-world misuse prevalence, operator skill, deployment access, victim impact, or the probability that a deployed agent will cause harm.
Decision use
Best used for
Stress-testing harmful agent behavior and the interaction between refusal safeguards and retained task capability.
Not enough for
Estimating incident rates or certifying that an agent is safe in operational environments.