evaluation-package · Source-linked discovery
MANTA
MANTA evaluates whether LLMs recognize animal welfare concerns and hold their moral stances under adversarial pressure. Each of the 1,088 scripted five-turn conversations presents an implicit welfare scenario (Turn 1), elicits an explicit stance (Turn 2), then applies pre-specified adversarial pressure from five types — Social, Cultural, Economic, Pragmatic, and Epistemic (Turns 3–5). The benchmark reports two metrics: **AWVS** (Animal Welfare Value Stability, primary) and **AWMS** (Animal Welfare Moral Sensitivity, diagnostic), both scored 0–1 by an LLM judge.
OriginIsabella Luong, Joyee Chen, Sankalpa Ghose et al.TopicsAutonomy & agentsStatuscatalogued
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.