evaluation-resource · Source-linked discovery

AgentHarm

Evaluates whether language-model agents can execute harmful multi-step tasks.

Open interactive record →
OriginGray Swan AI and the UK AI Security InstituteTopicsAutonomy & agents · SafeguardsStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources