evaluation-task · Source-linked discovery
InstrumentalEval - Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?
An evaluation designed to detect instrumental convergence behaviors in model responses (e.g., self-preservation, resource acquisition, power-seeking, strategic deception) using a rubric-driven LLM grader. The benchmark tests whether AI systems exhibit behaviors that are instrumentally useful across a wide range of goals, which could indicate concerning patterns of strategic reasoning.
OriginYufei He, Yuexin Li, Jiaying Wu et al.TopicsDeception & misalignmentStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.