evaluation-task · Source-linked discovery

ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation

Evaluates LLMs on class-level code generation with 100 tasks constructed over 500 person-hours. The study shows that LLMs perform worse on class-level tasks compared to method-level tasks.

Open interactive record →
OriginXueying Du, Mingwei Liu, Kaixin Wang et al.TopicsGeneral capabilityStatusimported

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources