evaluation-task · Source-linked discovery

MaCBench: Probing the limitations of multimodal language models for chemistry and materials research

MaCBench is a comprehensive benchmark for evaluating how vision-language models handle real-world chemistry and materials science tasks across three core aspects: data extraction, experimental understanding, and results interpretation. The dataset comprises over 1100 high-quality multimodal questions manually curated by chemistry and materials experts across 34 subsets.

Open interactive record →
OriginNawaf Alampara, Mara Schilling-Wilhelmi, Martiño Ríos-García et al.TopicsGeneral capabilityStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources