Chain-of-Descriptions: Improving Code LLMs for VHDL Code Generation and Summarization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Vijayaraghavan, Prashanth, Nitsure, Apoorva, Mackin, Charles, Shi, Luyao, Ambrogio, Stefano, Haran, Arvind, Paruthi, Viresh, Elzein, Ali, Coops, Dan, Beymer, David, Baldwin, Tyler, Degan, Ehsan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909691715518464
author Vijayaraghavan, Prashanth
Nitsure, Apoorva
Mackin, Charles
Shi, Luyao
Ambrogio, Stefano
Haran, Arvind
Paruthi, Viresh
Elzein, Ali
Coops, Dan
Beymer, David
Baldwin, Tyler
Degan, Ehsan
author_facet Vijayaraghavan, Prashanth
Nitsure, Apoorva
Mackin, Charles
Shi, Luyao
Ambrogio, Stefano
Haran, Arvind
Paruthi, Viresh
Elzein, Ali
Coops, Dan
Beymer, David
Baldwin, Tyler
Degan, Ehsan
contents Large Language Models (LLMs) have become widely used across diverse NLP tasks and domains, demonstrating their adaptability and effectiveness. In the realm of Electronic Design Automation (EDA), LLMs show promise for tasks like Register-Transfer Level (RTL) code generation and summarization. However, despite the proliferation of LLMs for general code-related tasks, there's a dearth of research focused on evaluating and refining these models for hardware description languages (HDLs), notably VHDL. In this study, we evaluate the performance of existing code LLMs for VHDL code generation and summarization using various metrics and two datasets -- VHDL-Eval and VHDL-Xform. The latter, an in-house dataset, aims to gauge LLMs' understanding of functionally equivalent code. Our findings reveal consistent underperformance of these models across different metrics, underscoring a significant gap in their suitability for this domain. To address this challenge, we propose Chain-of-Descriptions (CoDes), a novel approach to enhance the performance of LLMs for VHDL code generation and summarization tasks. CoDes involves generating a series of intermediate descriptive steps based on: (i) the problem statement for code generation, and (ii) the VHDL code for summarization. These steps are then integrated with the original input prompt (problem statement or code) and provided as input to the LLMs to generate the final output. Our experiments demonstrate that the CoDes approach significantly surpasses the standard prompting strategy across various metrics on both datasets. This method not only improves the quality of VHDL code generation and summarization but also serves as a framework for future research aimed at enhancing code LLMs for VHDL.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12308
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chain-of-Descriptions: Improving Code LLMs for VHDL Code Generation and Summarization
Vijayaraghavan, Prashanth
Nitsure, Apoorva
Mackin, Charles
Shi, Luyao
Ambrogio, Stefano
Haran, Arvind
Paruthi, Viresh
Elzein, Ali
Coops, Dan
Beymer, David
Baldwin, Tyler
Degan, Ehsan
Computation and Language
Artificial Intelligence
Hardware Architecture
Large Language Models (LLMs) have become widely used across diverse NLP tasks and domains, demonstrating their adaptability and effectiveness. In the realm of Electronic Design Automation (EDA), LLMs show promise for tasks like Register-Transfer Level (RTL) code generation and summarization. However, despite the proliferation of LLMs for general code-related tasks, there's a dearth of research focused on evaluating and refining these models for hardware description languages (HDLs), notably VHDL. In this study, we evaluate the performance of existing code LLMs for VHDL code generation and summarization using various metrics and two datasets -- VHDL-Eval and VHDL-Xform. The latter, an in-house dataset, aims to gauge LLMs' understanding of functionally equivalent code. Our findings reveal consistent underperformance of these models across different metrics, underscoring a significant gap in their suitability for this domain. To address this challenge, we propose Chain-of-Descriptions (CoDes), a novel approach to enhance the performance of LLMs for VHDL code generation and summarization tasks. CoDes involves generating a series of intermediate descriptive steps based on: (i) the problem statement for code generation, and (ii) the VHDL code for summarization. These steps are then integrated with the original input prompt (problem statement or code) and provided as input to the LLMs to generate the final output. Our experiments demonstrate that the CoDes approach significantly surpasses the standard prompting strategy across various metrics on both datasets. This method not only improves the quality of VHDL code generation and summarization but also serves as a framework for future research aimed at enhancing code LLMs for VHDL.
title Chain-of-Descriptions: Improving Code LLMs for VHDL Code Generation and Summarization
topic Computation and Language
Artificial Intelligence
Hardware Architecture
url https://arxiv.org/abs/2507.12308