LoRA Users Beware: A Few Spurious Tokens Can Manipulate Your Finetuned Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Salles, Marcel Mateos, Goyal, Praney, Sekhsaria, Pradyut, Huang, Hai, Balestriero, Randall
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915526914080768
author Salles, Marcel Mateos
Goyal, Praney
Sekhsaria, Pradyut
Huang, Hai
Balestriero, Randall
author_facet Salles, Marcel Mateos
Goyal, Praney
Sekhsaria, Pradyut
Huang, Hai
Balestriero, Randall
contents Large Language Models (LLMs) are commonly finetuned for a variety of use cases and domains. A common approach is to leverage Low-Rank Adaptation (LoRA) -- known to provide strong performance at low resource costs. In this study, we demonstrate that LoRA actually opens the door to short-cut vulnerabilities -- and the more resource efficient is the LoRA setup, the more vulnerable will be the finetuned model to aggressive attacks. To measure that vulnerability, we introduce Seamless Spurious Token Injection (SSTI), where we find that LoRA exclusively focuses on even just a single token that is spuriously correlated with downstream labels. In short, injection of that spurious token during finetuning ensure that the model's prediction at test-time can be manipulated on-demand. We conducted experiments across model families and datasets to evaluate the impact of SSTI during LoRA finetuning while providing possible mitigations. Our experiments conclude that none of the existing checkers and preprocessors can sanitize a dataset raising new concerns for data quality and AI safety.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11402
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LoRA Users Beware: A Few Spurious Tokens Can Manipulate Your Finetuned Model
Salles, Marcel Mateos
Goyal, Praney
Sekhsaria, Pradyut
Huang, Hai
Balestriero, Randall
Machine Learning
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) are commonly finetuned for a variety of use cases and domains. A common approach is to leverage Low-Rank Adaptation (LoRA) -- known to provide strong performance at low resource costs. In this study, we demonstrate that LoRA actually opens the door to short-cut vulnerabilities -- and the more resource efficient is the LoRA setup, the more vulnerable will be the finetuned model to aggressive attacks. To measure that vulnerability, we introduce Seamless Spurious Token Injection (SSTI), where we find that LoRA exclusively focuses on even just a single token that is spuriously correlated with downstream labels. In short, injection of that spurious token during finetuning ensure that the model's prediction at test-time can be manipulated on-demand. We conducted experiments across model families and datasets to evaluate the impact of SSTI during LoRA finetuning while providing possible mitigations. Our experiments conclude that none of the existing checkers and preprocessors can sanitize a dataset raising new concerns for data quality and AI safety.
title LoRA Users Beware: A Few Spurious Tokens Can Manipulate Your Finetuned Model
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.11402