Evaluation topic
Deception & misalignment
Scheming, covert action, strategic deception, sandbagging and misaligned agent behaviour.
20 records are tagged here. Inclusion does not establish construct equivalence, completeness, or comparability.
- Agentic Misalignment: How LLMs could be insider threatsAnthropic
- Alignment FakingRyan Greenblatt, Carson Denison, Benjamin Wright et al.
- DeceptionBenchPKU-Alignment and the DeceptionBench paper authors
- GDM Cover Your TracksGoogle DeepMind
- GDM Oversight PatternGoogle DeepMind
- GDM Self Reasoning Approved DirectoriesGoogle DeepMind
- GDM Self Reasoning Calculator ImprovementGoogle DeepMind
- GDM Self Reasoning Context Length Mod Instrumental OnlyGoogle DeepMind
- GDM Self Reasoning Context Length Mod Irreversibility OnlyGoogle DeepMind
- GDM Self Reasoning Database ToolGoogle DeepMind
- GDM Self Reasoning Latency CalculatorGoogle DeepMind
- GDM Self Reasoning Max Messages CalculatorGoogle DeepMind
- GDM Self Reasoning Max TokensGoogle DeepMind
- GDM Self Reasoning Oversight FrequencyGoogle DeepMind
- GDM Self Reasoning Read LogsGoogle DeepMind
- GDM Self Reasoning Turn Off FiltersGoogle DeepMind
- GDM Strategic Rule BreakingGoogle DeepMind
- InstrumentalEval - Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?Yufei He, Yuexin Li, Jiaying Wu et al.
- MonitorBench (Goal sandbag math subset)Han Wang, Yifan Sun, Brian Ko et al.
- Sabotage EvaluationsAnthropic