evaluation-package · Source-linked discovery

ArxivRollBench

A rolling benchmark for evaluating recent scientific text reasoning from arXiv papers. ArxivRollBench constructs multiple-choice sequencing, cloze, and next-fragment prediction tasks over newly released scientific papers across arXiv domains, with compact and full public releases.

Open interactive record →
OriginZi Liang, Liantong Yu, Shiyu Zhang et al.TopicsEvaluation integrityStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources