evaluation-task · Source-linked discovery
ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation
Evaluates LLMs on class-level code generation with 100 tasks constructed over 500 person-hours. The study shows that LLMs perform worse on class-level tasks compared to method-level tasks.
OriginXueying Du, Mingwei Liu, Kaixin Wang et al.TopicsGeneral capabilityStatusimported
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.