Trojans in Large Language Models of Code: A Critical Review through a Trigger-Based Taxonomy

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hussain, Aftab, Rabin, Md Rafiqul Islam, Ahmed, Toufique, Xu, Bowen, Devanbu, Premkumar, Alipour, Mohammad Amin
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917657295454208
author Hussain, Aftab
Rabin, Md Rafiqul Islam
Ahmed, Toufique
Xu, Bowen
Devanbu, Premkumar
Alipour, Mohammad Amin
author_facet Hussain, Aftab
Rabin, Md Rafiqul Islam
Ahmed, Toufique
Xu, Bowen
Devanbu, Premkumar
Alipour, Mohammad Amin
contents Large language models (LLMs) have provided a lot of exciting new capabilities in software development. However, the opaque nature of these models makes them difficult to reason about and inspect. Their opacity gives rise to potential security risks, as adversaries can train and deploy compromised models to disrupt the software development process in the victims' organization. This work presents an overview of the current state-of-the-art trojan attacks on large language models of code, with a focus on triggers -- the main design point of trojans -- with the aid of a novel unifying trigger taxonomy framework. We also aim to provide a uniform definition of the fundamental concepts in the area of trojans in Code LLMs. Finally, we draw implications of findings on how code models learn on trigger design.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02828
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Trojans in Large Language Models of Code: A Critical Review through a Trigger-Based Taxonomy
Hussain, Aftab
Rabin, Md Rafiqul Islam
Ahmed, Toufique
Xu, Bowen
Devanbu, Premkumar
Alipour, Mohammad Amin
Software Engineering
Machine Learning
Large language models (LLMs) have provided a lot of exciting new capabilities in software development. However, the opaque nature of these models makes them difficult to reason about and inspect. Their opacity gives rise to potential security risks, as adversaries can train and deploy compromised models to disrupt the software development process in the victims' organization. This work presents an overview of the current state-of-the-art trojan attacks on large language models of code, with a focus on triggers -- the main design point of trojans -- with the aid of a novel unifying trigger taxonomy framework. We also aim to provide a uniform definition of the fundamental concepts in the area of trojans in Code LLMs. Finally, we draw implications of findings on how code models learn on trigger design.
title Trojans in Large Language Models of Code: A Critical Review through a Trigger-Based Taxonomy
topic Software Engineering
Machine Learning
url https://arxiv.org/abs/2405.02828