evaluation-task · Source-linked discovery
MaCBench: Probing the limitations of multimodal language models for chemistry and materials research
MaCBench is a comprehensive benchmark for evaluating how vision-language models handle real-world chemistry and materials science tasks across three core aspects: data extraction, experimental understanding, and results interpretation. The dataset comprises over 1100 high-quality multimodal questions manually curated by chemistry and materials experts across 34 subsets.
OriginNawaf Alampara, Mara Schilling-Wilhelmi, Martiño Ríos-García et al.TopicsGeneral capabilityStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.