Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Zi, Hu, Haibo, Ye, Qingqing, Xiao, Yaxin, Li, Ronghua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912381989289984
author Liang, Zi
Hu, Haibo
Ye, Qingqing
Xiao, Yaxin
Li, Ronghua
author_facet Liang, Zi
Hu, Haibo
Ye, Qingqing
Xiao, Yaxin
Li, Ronghua
contents Low rank adaptation (LoRA) has emerged as a prominent technique for fine-tuning large language models (LLMs) thanks to its superb efficiency gains over previous methods. While extensive studies have examined the performance and structural properties of LoRA, its behavior upon training-time attacks remain underexplored, posing significant security risks. In this paper, we theoretically investigate the security implications of LoRA's low-rank structure during fine-tuning, in the context of its robustness against data poisoning and backdoor attacks. We propose an analytical framework that models LoRA's training dynamics, employs the neural tangent kernel to simplify the analysis of the training process, and applies information theory to establish connections between LoRA's low rank structure and its vulnerability against training-time attacks. Our analysis indicates that LoRA exhibits better robustness to backdoor attacks than full fine-tuning, while becomes more vulnerable to untargeted data poisoning due to its over-simplified information geometry. Extensive experimental evaluations have corroborated our theoretical findings.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12871
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?
Liang, Zi
Hu, Haibo
Ye, Qingqing
Xiao, Yaxin
Li, Ronghua
Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
Low rank adaptation (LoRA) has emerged as a prominent technique for fine-tuning large language models (LLMs) thanks to its superb efficiency gains over previous methods. While extensive studies have examined the performance and structural properties of LoRA, its behavior upon training-time attacks remain underexplored, posing significant security risks. In this paper, we theoretically investigate the security implications of LoRA's low-rank structure during fine-tuning, in the context of its robustness against data poisoning and backdoor attacks. We propose an analytical framework that models LoRA's training dynamics, employs the neural tangent kernel to simplify the analysis of the training process, and applies information theory to establish connections between LoRA's low rank structure and its vulnerability against training-time attacks. Our analysis indicates that LoRA exhibits better robustness to backdoor attacks than full fine-tuning, while becomes more vulnerable to untargeted data poisoning due to its over-simplified information geometry. Extensive experimental evaluations have corroborated our theoretical findings.
title Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?
topic Machine Learning
Artificial Intelligence
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2505.12871