RealBench: Benchmarking Verilog Generation Models with Real-World IP Designs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909698691694592 |
|---|---|
| author | Jin, Pengwei Huang, Di Li, Chongxiao Cheng, Shuyao Zhao, Yang Zheng, Xinyao Zhu, Jiaguo Xing, Shuyi Dou, Bohan Zhang, Rui Du, Zidong Guo, Qi Hu, Xing |
| author_facet | Jin, Pengwei Huang, Di Li, Chongxiao Cheng, Shuyao Zhao, Yang Zheng, Xinyao Zhu, Jiaguo Xing, Shuyi Dou, Bohan Zhang, Rui Du, Zidong Guo, Qi Hu, Xing |
| contents | The automatic generation of Verilog code using Large Language Models (LLMs) has garnered significant interest in hardware design automation. However, existing benchmarks for evaluating LLMs in Verilog generation fall short in replicating real-world design workflows due to their designs' simplicity, inadequate design specifications, and less rigorous verification environments. To address these limitations, we present RealBench, the first benchmark aiming at real-world IP-level Verilog generation tasks. RealBench features complex, structured, real-world open-source IP designs, multi-modal and formatted design specifications, and rigorous verification environments, including 100% line coverage testbenches and a formal checker. It supports both module-level and system-level tasks, enabling comprehensive assessments of LLM capabilities. Evaluations on various LLMs and agents reveal that even one of the best-performing LLMs, o1-preview, achieves only a 13.3% pass@1 on module-level tasks and 0% on system-level tasks, highlighting the need for stronger Verilog generation models in the future. The benchmark is open-sourced at https://github.com/IPRC-DIP/RealBench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_16200 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RealBench: Benchmarking Verilog Generation Models with Real-World IP Designs Jin, Pengwei Huang, Di Li, Chongxiao Cheng, Shuyao Zhao, Yang Zheng, Xinyao Zhu, Jiaguo Xing, Shuyi Dou, Bohan Zhang, Rui Du, Zidong Guo, Qi Hu, Xing Machine Learning Hardware Architecture The automatic generation of Verilog code using Large Language Models (LLMs) has garnered significant interest in hardware design automation. However, existing benchmarks for evaluating LLMs in Verilog generation fall short in replicating real-world design workflows due to their designs' simplicity, inadequate design specifications, and less rigorous verification environments. To address these limitations, we present RealBench, the first benchmark aiming at real-world IP-level Verilog generation tasks. RealBench features complex, structured, real-world open-source IP designs, multi-modal and formatted design specifications, and rigorous verification environments, including 100% line coverage testbenches and a formal checker. It supports both module-level and system-level tasks, enabling comprehensive assessments of LLM capabilities. Evaluations on various LLMs and agents reveal that even one of the best-performing LLMs, o1-preview, achieves only a 13.3% pass@1 on module-level tasks and 0% on system-level tasks, highlighting the need for stronger Verilog generation models in the future. The benchmark is open-sourced at https://github.com/IPRC-DIP/RealBench. |
| title | RealBench: Benchmarking Verilog Generation Models with Real-World IP Designs |
| topic | Machine Learning Hardware Architecture |
| url | https://arxiv.org/abs/2507.16200 |