Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Improta, Cristina
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914709434793984
author Improta, Cristina
author_facet Improta, Cristina
contents AI-based code generators have gained a fundamental role in assisting developers in writing software starting from natural language (NL). However, since these large language models are trained on massive volumes of data collected from unreliable online sources (e.g., GitHub, Hugging Face), AI models become an easy target for data poisoning attacks, in which an attacker corrupts the training data by injecting a small amount of poison into it, i.e., astutely crafted malicious samples. In this position paper, we address the security of AI code generators by identifying a novel data poisoning attack that results in the generation of vulnerable code. Next, we devise an extensive evaluation of how these attacks impact state-of-the-art models for code generation. Lastly, we discuss potential solutions to overcome this threat.
format Preprint
id arxiv_https___arxiv_org_abs_2403_06675
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
Improta, Cristina
Cryptography and Security
Artificial Intelligence
Software Engineering
AI-based code generators have gained a fundamental role in assisting developers in writing software starting from natural language (NL). However, since these large language models are trained on massive volumes of data collected from unreliable online sources (e.g., GitHub, Hugging Face), AI models become an easy target for data poisoning attacks, in which an attacker corrupts the training data by injecting a small amount of poison into it, i.e., astutely crafted malicious samples. In this position paper, we address the security of AI code generators by identifying a novel data poisoning attack that results in the generation of vulnerable code. Next, we devise an extensive evaluation of how these attacks impact state-of-the-art models for code generation. Lastly, we discuss potential solutions to overcome this threat.
title Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
topic Cryptography and Security
Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2403.06675