Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pinckney, Nathaniel, Deng, Chenhui, Ho, Chia-Tung, Tsai, Yun-Da, Liu, Mingjie, Zhou, Wenfei, Khailany, Brucek, Ren, Haoxing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916796651536384
author Pinckney, Nathaniel
Deng, Chenhui
Ho, Chia-Tung
Tsai, Yun-Da
Liu, Mingjie
Zhou, Wenfei
Khailany, Brucek
Ren, Haoxing
author_facet Pinckney, Nathaniel
Deng, Chenhui
Ho, Chia-Tung
Tsai, Yun-Da
Liu, Mingjie
Zhou, Wenfei
Khailany, Brucek
Ren, Haoxing
contents We present the Comprehensive Verilog Design Problems (CVDP) benchmark, a new dataset and infrastructure to advance LLM and agent research in hardware design and verification. CVDP includes 783 problems across 13 task categories, covering RTL generation, verification, debugging, specification alignment, and technical Q&A authored by experienced hardware engineers. Problems are offered in both non-agentic and agentic formats. The benchmark introduces more realistic and challenging contexts than prior work, with state-of-the-art models achieving no more than 34% pass@1 on code generation. Agentic tasks$\unicode{x2013}$especially those involving RTL reuse and verification$\unicode{x2013}$are particularly difficult. Evaluation uses open-source tools and model scoring infrastructure, with comprehension tasks assessed via BLEU and LLM-based judging. CVDP reveals substantial gaps in current model capabilities, underscoring the need for continued research toward robust, real-world hardware design automation.
format Preprint
id arxiv_https___arxiv_org_abs_2506_14074
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification
Pinckney, Nathaniel
Deng, Chenhui
Ho, Chia-Tung
Tsai, Yun-Da
Liu, Mingjie
Zhou, Wenfei
Khailany, Brucek
Ren, Haoxing
Machine Learning
Hardware Architecture
We present the Comprehensive Verilog Design Problems (CVDP) benchmark, a new dataset and infrastructure to advance LLM and agent research in hardware design and verification. CVDP includes 783 problems across 13 task categories, covering RTL generation, verification, debugging, specification alignment, and technical Q&A authored by experienced hardware engineers. Problems are offered in both non-agentic and agentic formats. The benchmark introduces more realistic and challenging contexts than prior work, with state-of-the-art models achieving no more than 34% pass@1 on code generation. Agentic tasks$\unicode{x2013}$especially those involving RTL reuse and verification$\unicode{x2013}$are particularly difficult. Evaluation uses open-source tools and model scoring infrastructure, with comprehension tasks assessed via BLEU and LLM-based judging. CVDP reveals substantial gaps in current model capabilities, underscoring the need for continued research toward robust, real-world hardware design automation.
title Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification
topic Machine Learning
Hardware Architecture
url https://arxiv.org/abs/2506.14074