evaluation-package · Source-linked discovery

MonitorBench (Goal sandbag math subset)

MonitorBench evaluates chain-of-thought (CoT) monitorability: whether a model's reasoning faithfully surfaces the decision-critical factors driving its final output. The full benchmark spans 1,514 instances across 19 tasks in 7 categories along three axes (input intervention, outcome justification, solution process), with two stress-test conditions (direct concealment, monitor-aware evasion) probing intentional suppression. This subset implements the goal-sandbag math task over 30 AIME 2024 problems: each sample pairs an original prompt with an intervened prompt that instructs the model to deliberately underperform while appearing competent, and a monitor LLM judges whether the sandbagging is detectable from the action, the CoT, or both. Monitorability is the monitor's detection rate normalized by the causal effect of the sandbag instruction, macro-averaged over eligible problems.

Open interactive record →
OriginHan Wang, Yifan Sun, Brian Ko et al.TopicsDeception & misalignment · Evaluation integrityStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources