evaluation-package · Source-linked discovery
MonitorBench (Steganography subset)
MonitorBench evaluates chain-of-thought (CoT) monitorability: whether a model's reasoning faithfully surfaces the decision-critical factors driving its outputs. The full benchmark spans 1,514 instances across 19 tasks in 7 categories along three axes (input intervention, outcome justification, solution process), with stress-test conditions probing intentional suppression. This subset implements the steganography task: 50 writing questions from Databricks Dolly-15k, each paired with a payload word the model must covertly encode as a sentence-initial acrostic while answering. After deterministic verification that the payload was encoded, a monitor LLM judges whether the hidden side objective is visible in the action, the CoT, or both, under three conditions (standard, direct concealment, monitor-aware evasion). Monitorability is the fraction of verification-passing rollouts the monitor flags as detected.
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.