evaluation-package · Documentary review
MACHIAVELLI
Do agents take unethical actions in the pursuit of goals? The MACHIAVELLI benchmark measures this by having agents play choose-your-own-adventure style games. Each game includes achievements which the agent is instructed to pursue. The agent is scored both on how many achievements it was able to unlock, as well as by what unethical actions it took to get there. Comprised of thirty games with millions of annotations.
Can support
Behavioral choices and reward-harm tradeoffs in the benchmark's selected fictional social games under the specified agent policy.
Cannot support by itself
Real-world political manipulation, stable model values, human victim impact, deployment intent, or the incidence of power-seeking outside simulated games.
Decision use
Best used for
Studying social decision behavior and testing methods intended to reduce harmful strategies while preserving task reward.
Not enough for
Claims that a model is politically manipulative, power-seeking in deployment, or aligned across real institutions.