VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wei, Anjiang, Tan, Huanmi, Suresh, Tarun, Mendoza, Daniel, Teixeira, Thiago S. F. X., Wang, Ke, Trippel, Caroline, Aiken, Alex
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918129400020992
author Wei, Anjiang
Tan, Huanmi
Suresh, Tarun
Mendoza, Daniel
Teixeira, Thiago S. F. X.
Wang, Ke
Trippel, Caroline
Aiken, Alex
author_facet Wei, Anjiang
Tan, Huanmi
Suresh, Tarun
Mendoza, Daniel
Teixeira, Thiago S. F. X.
Wang, Ke
Trippel, Caroline
Aiken, Alex
contents Recent advances in Large Language Models (LLMs) have sparked growing interest in applying them to Electronic Design Automation (EDA) tasks, particularly Register Transfer Level (RTL) code generation. While several RTL datasets have been introduced, most focus on syntactic validity rather than functional validation with tests, leading to training examples that compile but may not implement the intended behavior. We present VERICODER, a model for RTL code generation fine-tuned on a dataset validated for functional correctness. This fine-tuning dataset is constructed using a novel methodology that combines unit test generation with feedback-directed refinement. Given a natural language specification and an initial RTL design, we prompt a teacher model (GPT-4o-mini) to generate unit tests and iteratively revise the RTL design based on its simulation results using the generated tests. If necessary, the teacher model also updates the tests to ensure they comply with the natural language specification. As a result of this process, every example in our dataset is functionally validated, consisting of a natural language description, an RTL implementation, and passing tests. Fine-tuned on this dataset of 125,777 examples, VERICODER achieves state-of-the-art metrics in functional correctness on VerilogEval and RTLLM, with relative gains of up to 71.7% and 27.4%, respectively. An ablation study further shows that models trained on our functionally validated dataset outperform those trained on functionally non-validated datasets, underscoring the importance of high-quality datasets in RTL code generation. Our code, data, and models are publicly available at https://github.com/Anjiang-Wei/VeriCoder
format Preprint
id arxiv_https___arxiv_org_abs_2504_15659
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation
Wei, Anjiang
Tan, Huanmi
Suresh, Tarun
Mendoza, Daniel
Teixeira, Thiago S. F. X.
Wang, Ke
Trippel, Caroline
Aiken, Alex
Hardware Architecture
Artificial Intelligence
Computation and Language
Machine Learning
Software Engineering
Recent advances in Large Language Models (LLMs) have sparked growing interest in applying them to Electronic Design Automation (EDA) tasks, particularly Register Transfer Level (RTL) code generation. While several RTL datasets have been introduced, most focus on syntactic validity rather than functional validation with tests, leading to training examples that compile but may not implement the intended behavior. We present VERICODER, a model for RTL code generation fine-tuned on a dataset validated for functional correctness. This fine-tuning dataset is constructed using a novel methodology that combines unit test generation with feedback-directed refinement. Given a natural language specification and an initial RTL design, we prompt a teacher model (GPT-4o-mini) to generate unit tests and iteratively revise the RTL design based on its simulation results using the generated tests. If necessary, the teacher model also updates the tests to ensure they comply with the natural language specification. As a result of this process, every example in our dataset is functionally validated, consisting of a natural language description, an RTL implementation, and passing tests. Fine-tuned on this dataset of 125,777 examples, VERICODER achieves state-of-the-art metrics in functional correctness on VerilogEval and RTLLM, with relative gains of up to 71.7% and 27.4%, respectively. An ablation study further shows that models trained on our functionally validated dataset outperform those trained on functionally non-validated datasets, underscoring the importance of high-quality datasets in RTL code generation. Our code, data, and models are publicly available at https://github.com/Anjiang-Wei/VeriCoder
title VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation
topic Hardware Architecture
Artificial Intelligence
Computation and Language
Machine Learning
Software Engineering
url https://arxiv.org/abs/2504.15659