Saved in:
Bibliographic Details
Main Authors: Luong, Hoang-Chau, Chen, Lingwei
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2601.06305
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909986523709440
author Luong, Hoang-Chau
Chen, Lingwei
author_facet Luong, Hoang-Chau
Chen, Lingwei
contents Low-Rank Adaptation (LoRA) is widely used for parameter-efficient fine-tuning of large language models, but it is notably ineffective at removing backdoor behaviors from poisoned pretrained models when fine-tuning on clean dataset. Contrary to the common belief that this weakness is caused primarily by low rank, we show that LoRA's vulnerability is fundamentally spectral. Our analysis identifies two key factors: LoRA updates (i) possess insufficient spectral strength, with singular values far below those of pretrained weights, and (ii) exhibit unfavorable spectral alignment, weakly matching clean-task directions while retaining overlap with trigger-sensitive subspaces. We further establish a critical scaling threshold beyond which LoRA can theoretically suppress trigger-induced activations, and we show empirically that standard LoRA rarely reaches this regime. We introduce Regularized Low-Rank Adaptation (RoRA), which improves forgetting by increasing spectral strength and correcting alignment through clean-strengthened regularization, trigger-insensitive constraints, and post-training spectral rescaling. Experiments across multiple NLP benchmarks and attack settings show that RoRA substantially reduces attack success rates while maintaining clean accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06305
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Why LoRA Fails to Forget: Regularized Low-Rank Adaptation Against Backdoors in Language Models
Luong, Hoang-Chau
Chen, Lingwei
Computation and Language
Low-Rank Adaptation (LoRA) is widely used for parameter-efficient fine-tuning of large language models, but it is notably ineffective at removing backdoor behaviors from poisoned pretrained models when fine-tuning on clean dataset. Contrary to the common belief that this weakness is caused primarily by low rank, we show that LoRA's vulnerability is fundamentally spectral. Our analysis identifies two key factors: LoRA updates (i) possess insufficient spectral strength, with singular values far below those of pretrained weights, and (ii) exhibit unfavorable spectral alignment, weakly matching clean-task directions while retaining overlap with trigger-sensitive subspaces. We further establish a critical scaling threshold beyond which LoRA can theoretically suppress trigger-induced activations, and we show empirically that standard LoRA rarely reaches this regime. We introduce Regularized Low-Rank Adaptation (RoRA), which improves forgetting by increasing spectral strength and correcting alignment through clean-strengthened regularization, trigger-insensitive constraints, and post-training spectral rescaling. Experiments across multiple NLP benchmarks and attack settings show that RoRA substantially reduces attack success rates while maintaining clean accuracy.
title Why LoRA Fails to Forget: Regularized Low-Rank Adaptation Against Backdoors in Language Models
topic Computation and Language
url https://arxiv.org/abs/2601.06305