evaluation-package · Source-linked discovery

MedCalc-Bench

Evaluates whether language models can perform clinical calculations from free-text patient notes. Each sample provides a patient note and asks the model to compute a specific clinical value across 55 calculators (lab tests, physical quantities, risk/severity scores, dosages, and dates). Rule-based calculators are scored by exact match; equation-based calculators are scored within the authors' tolerance band. Faithful Inspect AI port of the NeurIPS 2024 benchmark.

Open interactive record →
OriginNikhil Khandekar, Qiao Jin, Guangzhi Xiong et al.TopicsGeneral capabilityStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources