evaluation-task · Source-linked discovery

ChemBench: Are large language models superhuman chemists?

ChemBench is designed to reveal limitations of current frontier models for use in the chemical sciences. It consists of 2786 question-answer pairs compiled from diverse sources. Our corpus measures reasoning, knowledge and intuition across a large fraction of the topics taught in undergraduate and graduate chemistry curricula. It can be used to evaluate any system that can return text (i.e., including tool-augmented systems).

Open interactive record →
OriginAdrian Mirza, Nawaf Alampara, Sreekanth Kunchapu et al.TopicsBio / CBRNStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources