evaluation-task · Source-linked discovery
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
This benchmark evaluates LLM-based research agents on their ability to propose and implement novel methods using tasks from recent ML conference competitions, assessing both novelty and effectiveness compared to a baseline and top human solutions.
OriginYunxiang Zhang, Muhammad Khalifa, Shitanshu Bhushan et al.TopicsAI R&DStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.