evaluation-package · Source-linked discovery

ChipBench — Reference Model

ChipBench evaluates LLMs on chip design tasks using 264 test cases: 44 Verilog generation problems (self-contained, hierarchical, CPU IP modules), 89 debugging cases (arithmetic, assignment, timing, state-machine bugs, zero/one-shot), and 132 reference model generation samples (Python, SystemC, CXXRTL). Scoring uses iVerilog functional simulation against golden implementations with 1000+ random stimuli. Metrics are pass@1/5/10 across 20 samples per problem at temperature=0.85, top_p=0.95.

Open interactive record →
OriginZhongkai Yu, Chenyang Zhou, Yichen Lin et al.TopicsGeneral capabilityStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources