evaluation-task · Source-linked discovery

MBPP: Basic Python Coding Challenges

Measures the ability of language models to generate short Python programs from simple natural-language descriptions, testing basic coding proficiency.

Open interactive record →
OriginJacob Austin, Augustus Odena, Maxwell Nye et al.TopicsGeneral capabilityStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources