evaluation-task · Source-linked discovery
HumanEval: Python Function Generation from Instructions
Assesses how accurately language models can write correct Python functions based solely on natural-language instructions provided as docstrings.
OriginMark Chen, Jerry Tworek, Heewoo Jun et al.TopicsGeneral capabilityStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.