evaluation-package · Source-linked discovery
ChipBench — Debugging
ChipBench evaluates LLMs on AI-aided chip design across three tasks: Verilog generation (44 modules), Verilog debugging (89 cases), and reference model generation (132 cases in Python/SystemC/CXXRTL). The debugging task injects four bug types (arithmetic, assignment, timing, state machine) into golden Verilog modules; models receive buggy code plus a description (zero-shot) or additionally a VCD waveform (one-shot). Scoring uses iVerilog simulation against a golden reference, reporting pass@1/5/10 over 20 samples.
OriginZhongkai Yu, Chenyang Zhou, Yichen Lin et al.TopicsGeneral capabilityStatuscatalogued
Can support
Not independently assessed by FronteraEval yet.
Cannot support by itself
No inference beyond the upstream source should be made until the protocol is reviewed.