VeriContaminated: Assessing LLM-Driven Verilog Coding for Data Contamination

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Zeng, Shao, Minghao, Bhandari, Jitendra, Mankali, Likhitha, Karri, Ramesh, Sinanoglu, Ozgur, Shafique, Muhammad, Knechtel, Johann
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913889908686848
author Wang, Zeng
Shao, Minghao
Bhandari, Jitendra
Mankali, Likhitha
Karri, Ramesh
Sinanoglu, Ozgur
Shafique, Muhammad
Knechtel, Johann
author_facet Wang, Zeng
Shao, Minghao
Bhandari, Jitendra
Mankali, Likhitha
Karri, Ramesh
Sinanoglu, Ozgur
Shafique, Muhammad
Knechtel, Johann
contents Large Language Models (LLMs) have revolutionized code generation, achieving exceptional results on various established benchmarking frameworks. However, concerns about data contamination - where benchmark data inadvertently leaks into pre-training or fine-tuning datasets - raise questions about the validity of these evaluations. While this issue is known, limiting the industrial adoption of LLM-driven software engineering, hardware coding has received little to no attention regarding these risks. For the first time, we analyze state-of-the-art (SOTA) evaluation frameworks for Verilog code generation (VerilogEval and RTLLM), using established methods for contamination detection (CCD and Min-K% Prob). We cover SOTA commercial and open-source LLMs (CodeGen2.5, Minitron 4b, Mistral 7b, phi-4 mini, LLaMA-{1,2,3.1}, GPT-{2,3.5,4o}, Deepseek-Coder, and CodeQwen 1.5), in baseline and fine-tuned models (RTLCoder and Verigen). Our study confirms that data contamination is a critical concern. We explore mitigations and the resulting trade-offs for code quality vs fairness (i.e., reducing contamination toward unbiased benchmarking).
format Preprint
id arxiv_https___arxiv_org_abs_2503_13572
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VeriContaminated: Assessing LLM-Driven Verilog Coding for Data Contamination
Wang, Zeng
Shao, Minghao
Bhandari, Jitendra
Mankali, Likhitha
Karri, Ramesh
Sinanoglu, Ozgur
Shafique, Muhammad
Knechtel, Johann
Hardware Architecture
Cryptography and Security
Machine Learning
Large Language Models (LLMs) have revolutionized code generation, achieving exceptional results on various established benchmarking frameworks. However, concerns about data contamination - where benchmark data inadvertently leaks into pre-training or fine-tuning datasets - raise questions about the validity of these evaluations. While this issue is known, limiting the industrial adoption of LLM-driven software engineering, hardware coding has received little to no attention regarding these risks. For the first time, we analyze state-of-the-art (SOTA) evaluation frameworks for Verilog code generation (VerilogEval and RTLLM), using established methods for contamination detection (CCD and Min-K% Prob). We cover SOTA commercial and open-source LLMs (CodeGen2.5, Minitron 4b, Mistral 7b, phi-4 mini, LLaMA-{1,2,3.1}, GPT-{2,3.5,4o}, Deepseek-Coder, and CodeQwen 1.5), in baseline and fine-tuned models (RTLCoder and Verigen). Our study confirms that data contamination is a critical concern. We explore mitigations and the resulting trade-offs for code quality vs fairness (i.e., reducing contamination toward unbiased benchmarking).
title VeriContaminated: Assessing LLM-Driven Verilog Coding for Data Contamination
topic Hardware Architecture
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2503.13572