Inducing Vulnerable Code Generation in LLM Coding Assistants

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zeng, Binqi, Zhang, Quan, Zhou, Chijin, Go, Gwihwan, Jiang, Yu, Shi, Heyuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913803153702912
author Zeng, Binqi
Zhang, Quan
Zhou, Chijin
Go, Gwihwan
Jiang, Yu
Shi, Heyuan
author_facet Zeng, Binqi
Zhang, Quan
Zhou, Chijin
Go, Gwihwan
Jiang, Yu
Shi, Heyuan
contents Due to insufficient domain knowledge, LLM coding assistants often reference related solutions from the Internet to address programming problems. However, incorporating external information into LLMs' code generation process introduces new security risks. In this paper, we reveal a real-world threat, named HACKODE, where attackers exploit referenced external information to embed attack sequences, causing LLMs to produce code with vulnerabilities such as buffer overflows and incomplete validations. We designed a prototype of the attack, which generates effective attack sequences for potential diverse inputs with various user queries and prompt templates. Through the evaluation on two general LLMs and two code LLMs, we demonstrate that the attack is effective, achieving an 84.29% success rate. Additionally, on a real-world application, HACKODE achieves 75.92% ASR, demonstrating its real-world impact.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15867
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Inducing Vulnerable Code Generation in LLM Coding Assistants
Zeng, Binqi
Zhang, Quan
Zhou, Chijin
Go, Gwihwan
Jiang, Yu
Shi, Heyuan
Software Engineering
Due to insufficient domain knowledge, LLM coding assistants often reference related solutions from the Internet to address programming problems. However, incorporating external information into LLMs' code generation process introduces new security risks. In this paper, we reveal a real-world threat, named HACKODE, where attackers exploit referenced external information to embed attack sequences, causing LLMs to produce code with vulnerabilities such as buffer overflows and incomplete validations. We designed a prototype of the attack, which generates effective attack sequences for potential diverse inputs with various user queries and prompt templates. Through the evaluation on two general LLMs and two code LLMs, we demonstrate that the attack is effective, achieving an 84.29% success rate. Additionally, on a real-world application, HACKODE achieves 75.92% ASR, demonstrating its real-world impact.
title Inducing Vulnerable Code Generation in LLM Coding Assistants
topic Software Engineering
url https://arxiv.org/abs/2504.15867