evaluation-family · Source-linked discovery

Healthbench Meta Eval

A comprehensive evaluation benchmark designed to assess language models' medical capabilities across a wide range of healthcare scenarios.

Open interactive record →
OriginOpenAITopicsGeneral capabilityStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources