evaluation-package · Source-linked discovery

NARCBench

NARCBench evaluates detection of covert multi-agent collusion using public committee-deliberation transcripts. This Inspect wrapper presents a black-box LLM monitor with only public messages, final statements, and votes from NARCBench-Core runs (qwen3_32b or gpt_oss_20b backbones), asking it to output P(collusion) per run. Primary metric is AUROC of predicted probabilities against ground-truth collusion/control labels. The paper's own results use white-box activation probes; this eval measures text-level monitor performance on the same scenarios.

Open interactive record →
OriginAaron Rose, Carissa Cullen, Sahar Abdelnabi et al.TopicsSafeguardsStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources