SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Rezakhani, Mahshid, Mashnoor, Nowfel, Azar, Kimia, Kamali, Hadi
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913075331858432
author Rezakhani, Mahshid
Mashnoor, Nowfel
Azar, Kimia
Kamali, Hadi
author_facet Rezakhani, Mahshid
Mashnoor, Nowfel
Azar, Kimia
Kamali, Hadi
contents As large language models (LLMs) are increasingly fine-tuned for hardware tasks like RTL code generation, the scarcity of high-quality datasets often leads to the use of rapidly assembled or generated training data. These datasets frequently lack security verification and are highly susceptible to data poisoning attacks. Such poisoning can cause models to generate syntactically valid but insecure hardware modules that bypass standard functionality checks. To address this, we present SafeTune, a framework designed to harden LLM-based RTL generation against poisoning, specifically focusing on hardware Trojan (HT) insertion. SafeTune integrates two core components: (i) a Graph Neural Network (GNN) that models structural properties to identify anomalous circuitry patterns during fine-tuning, and (ii) a semantic verification module using text embeddings and an XGBoost classifier to assess prompt security. By coupling structural and semantic knowledge, SafeTune effectively filters poisoned inputs without sacrificing legitimate data. Experimental results demonstrate that SafeTune significantly enhances the robustness and reliability of LLM fine-tuning without requiring modifications to the underlying model architecture.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27238
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation
Rezakhani, Mahshid
Mashnoor, Nowfel
Azar, Kimia
Kamali, Hadi
Cryptography and Security
Hardware Architecture
As large language models (LLMs) are increasingly fine-tuned for hardware tasks like RTL code generation, the scarcity of high-quality datasets often leads to the use of rapidly assembled or generated training data. These datasets frequently lack security verification and are highly susceptible to data poisoning attacks. Such poisoning can cause models to generate syntactically valid but insecure hardware modules that bypass standard functionality checks. To address this, we present SafeTune, a framework designed to harden LLM-based RTL generation against poisoning, specifically focusing on hardware Trojan (HT) insertion. SafeTune integrates two core components: (i) a Graph Neural Network (GNN) that models structural properties to identify anomalous circuitry patterns during fine-tuning, and (ii) a semantic verification module using text embeddings and an XGBoost classifier to assess prompt security. By coupling structural and semantic knowledge, SafeTune effectively filters poisoned inputs without sacrificing legitimate data. Experimental results demonstrate that SafeTune significantly enhances the robustness and reliability of LLM fine-tuning without requiring modifications to the underlying model architecture.
title SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation
topic Cryptography and Security
Hardware Architecture
url https://arxiv.org/abs/2604.27238