evaluation-package · Source-linked discovery

ChipBench — Verilog Generation

ChipBench evaluates LLMs on Verilog code generation from natural-language specifications across three module categories: self-contained (29 cases), non-self-contained hierarchical (6 cases), and CPU IP submodules (9 cases). Each case provides a prompt, golden Verilog, and a test harness combining directed corner cases with 1000+ random stimuli compiled via iVerilog. Scoring uses pass@1/5/10 over 20 samples at temperature=0.85, top_p=0.95, executed in a Docker sandbox.

Open interactive record →
OriginZhongkai Yu, Chenyang Zhou, Yichen Lin et al.TopicsGeneral capabilityStatuscatalogued

Can support

Not independently assessed by FronteraEval yet.

Cannot support by itself

No inference beyond the upstream source should be made until the protocol is reviewed.

Original sources