evaluation-package · Source-linked discovery
MedCalc-Bench
Evaluates whether language models can perform clinical calculations from free-text patient notes. Each sample provides a patient note and asks the model to compute a specific clinical value across 55 calculators (lab tests, physical quantities, risk/severity scores, dosages, and dates). Rule-based calculators are scored by exact match; equation-based calculators are scored within the authors' tolerance band. Faithful Inspect AI port of the NeurIPS 2024 benchmark.
OriginNikhil Khandekar, Qiao Jin, Guangzhi Xiong et al.TopicsGeneral capabilityStatuscatalogued
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.