evaluation-task · Source-linked discovery
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
Evaluating models on multistep soft reasoning tasks in the form of free text narratives.
OriginZayne Sprague, Xi Ye, Kaj Bostrom et al.TopicsGeneral capabilityStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.