VerilogDB: The Largest, Highest-Quality Dataset with a Preprocessing Framework for LLM-based RTL Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Calzada, Paul E., Ibnat, Zahin, Rahman, Tanvir, Kandula, Kamal, Lu, Danyu, Saha, Sujan Kumar, Farahmandi, Farimah, Tehranipoor, Mark
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911063121854464
author Calzada, Paul E.
Ibnat, Zahin
Rahman, Tanvir
Kandula, Kamal
Lu, Danyu
Saha, Sujan Kumar
Farahmandi, Farimah
Tehranipoor, Mark
author_facet Calzada, Paul E.
Ibnat, Zahin
Rahman, Tanvir
Kandula, Kamal
Lu, Danyu
Saha, Sujan Kumar
Farahmandi, Farimah
Tehranipoor, Mark
contents Large Language Models (LLMs) are gaining popularity for hardware design automation, particularly through Register Transfer Level (RTL) code generation. In this work, we examine the current literature on RTL generation using LLMs and identify key requirements for training and fine-tuning datasets. We construct a robust Verilog dataset through an automated three-pronged process involving database (DB) creation and management with PostgreSQL, data collection from code hosting sites like OpenCores and GitHub, and data preprocessing to verify the codes' syntax, run logic synthesis, and extract relevant module metadata. We implement a scalable and efficient DB infrastructure to support analysis and detail our preprocessing pipeline to enforce high-quality data before DB insertion. The resulting dataset comprises 20,392 Verilog samples, 751 MB of Verilog code data, which is the largest high-quality Verilog dataset for LLM fine-tuning to our knowledge. We further evaluate the dataset, address associated challenges, and explore potential applications for future research and development in LLM-based hardware generation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13369
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VerilogDB: The Largest, Highest-Quality Dataset with a Preprocessing Framework for LLM-based RTL Generation
Calzada, Paul E.
Ibnat, Zahin
Rahman, Tanvir
Kandula, Kamal
Lu, Danyu
Saha, Sujan Kumar
Farahmandi, Farimah
Tehranipoor, Mark
Hardware Architecture
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) are gaining popularity for hardware design automation, particularly through Register Transfer Level (RTL) code generation. In this work, we examine the current literature on RTL generation using LLMs and identify key requirements for training and fine-tuning datasets. We construct a robust Verilog dataset through an automated three-pronged process involving database (DB) creation and management with PostgreSQL, data collection from code hosting sites like OpenCores and GitHub, and data preprocessing to verify the codes' syntax, run logic synthesis, and extract relevant module metadata. We implement a scalable and efficient DB infrastructure to support analysis and detail our preprocessing pipeline to enforce high-quality data before DB insertion. The resulting dataset comprises 20,392 Verilog samples, 751 MB of Verilog code data, which is the largest high-quality Verilog dataset for LLM fine-tuning to our knowledge. We further evaluate the dataset, address associated challenges, and explore potential applications for future research and development in LLM-based hardware generation.
title VerilogDB: The Largest, Highest-Quality Dataset with a Preprocessing Framework for LLM-based RTL Generation
topic Hardware Architecture
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.13369