evaluation-package · Source-linked discovery
ArxivRollBench
A rolling benchmark for evaluating recent scientific text reasoning from arXiv papers. ArxivRollBench constructs multiple-choice sequencing, cloze, and next-fragment prediction tasks over newly released scientific papers across arXiv domains, with compact and full public releases.
OriginZi Liang, Liantong Yu, Shiyu Zhang et al.TopicsEvaluation integrityStatuscatalogued
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.