evaluation-task · Source-linked discovery

MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?

This benchmark evaluates LLM-based research agents on their ability to propose and implement novel methods using tasks from recent ML conference competitions, assessing both novelty and effectiveness compared to a baseline and top human solutions.

Open interactive record →
OriginYunxiang Zhang, Muhammad Khalifa, Shitanshu Bhushan et al.TopicsAI R&DStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources