evaluation-task · Source-linked discovery
MBPP: Basic Python Coding Challenges
Measures the ability of language models to generate short Python programs from simple natural-language descriptions, testing basic coding proficiency.
OriginJacob Austin, Augustus Odena, Maxwell Nye et al.TopicsGeneral capabilityStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.