Evaluation topic
Evaluation integrity
Validity, contamination, elicitation, judge reliability, eval awareness and protocol integrity.
17 records are tagged here. Inclusion does not establish construct equivalence, completeness, or comparability.
- AbstentionBench: Reasoning LLMs Fail on Unanswerable QuestionsPolina Kirichenko, Mark Ibrahim, Kamalika Chaudhuri et al.
- AgentBoardHKUST NLP and the AgentBoard paper authors
- ARC-AGI-2ARC Prize Foundation and the ARC-AGI-2 authors
- ArxivRollBenchZi Liang, Liantong Yu, Shiyu Zhang et al.
- HELM SafetyStanford CRFM
- Inspect AIUK AI Security Institute
- JailbreakBenchJailbreakBench project and paper authors
- LiveBench: A Challenging, Contamination-Free LLM BenchmarkColin White, Samuel Dooley, Manley Roberts et al.
- lm-evaluation-harnessEleutherAI
- MonitorBench (Goal sandbag math subset)Han Wang, Yifan Sun, Brian Ko et al.
- MonitorBench (Steganography subset)Han Wang, Yifan Sun, Brian Ko et al.
- OpenAI EvalsOpenAI
- Persistbench JudgeSidharth Pulipaka, Oliver Chen, Manas Sharma et al.
- Pre-Flight: Aviation Operations Knowledge EvaluationAlex Brooker, Tim Hughes
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain Evans
- ZerobenchJonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma et al.
- Zerobench SubquestionsJonathan Roberts, Mohammad Reza Taesiri, Ansh Sharma et al.