Poison with Style: A Practical Poisoning Attack on Code Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tran, Khang, Boshmaf, Yazan, Khalil, Issa, Phan, NhatHai, Yu, Ting, Parvez, Md Rizwan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910263772446720
author Tran, Khang
Boshmaf, Yazan
Khalil, Issa
Phan, NhatHai
Yu, Ting
Parvez, Md Rizwan
author_facet Tran, Khang
Boshmaf, Yazan
Khalil, Issa
Phan, NhatHai
Yu, Ting
Parvez, Md Rizwan
contents Code Large Language Models (CLLMs) serve as the core of modern code agents, enabling developers to automate complex software development tasks. In this paper, we present Poison-with-Style (PwS), a practical and stealthy model poisoning attack targeting CLLMs. Unlike prior attacks that assume an active adversary capable of directly embedding explicit triggers (e.g., specific words) into developers' prompts during inference, PwS leverages developers' code styles as covert triggers implicitly embedded within their prompts. PwS introduces a novel data collection method and a two-step training strategy to fine-tune CLLMs, causing them to generate vulnerable code when prompts contain trigger code styles while maintaining normal behavior on other prompts. Experimental results on Python code completion tasks show that PwS is robust against state-of-the-art defenses and achieves high attack success rates across diverse vulnerabilities, while maintaining strong performance on standard code completion benchmarks. For example, PwS-poisoned models generate CWE-20 vulnerable code in 95% of cases when the trigger code style is used, with less than a 5% drop in pass@1 performance on the HumanEval and MBPP benchmarks. Our implementation and dataset are here: https://github.com/khangtran2020/pws.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27631
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Poison with Style: A Practical Poisoning Attack on Code Large Language Models
Tran, Khang
Boshmaf, Yazan
Khalil, Issa
Phan, NhatHai
Yu, Ting
Parvez, Md Rizwan
Cryptography and Security
Machine Learning
Code Large Language Models (CLLMs) serve as the core of modern code agents, enabling developers to automate complex software development tasks. In this paper, we present Poison-with-Style (PwS), a practical and stealthy model poisoning attack targeting CLLMs. Unlike prior attacks that assume an active adversary capable of directly embedding explicit triggers (e.g., specific words) into developers' prompts during inference, PwS leverages developers' code styles as covert triggers implicitly embedded within their prompts. PwS introduces a novel data collection method and a two-step training strategy to fine-tune CLLMs, causing them to generate vulnerable code when prompts contain trigger code styles while maintaining normal behavior on other prompts. Experimental results on Python code completion tasks show that PwS is robust against state-of-the-art defenses and achieves high attack success rates across diverse vulnerabilities, while maintaining strong performance on standard code completion benchmarks. For example, PwS-poisoned models generate CWE-20 vulnerable code in 95% of cases when the trigger code style is used, with less than a 5% drop in pass@1 performance on the HumanEval and MBPP benchmarks. Our implementation and dataset are here: https://github.com/khangtran2020/pws.
title Poison with Style: A Practical Poisoning Attack on Code Large Language Models
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2605.27631