Weight space Detection of Backdoors in LoRA Adapters

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Merenciano, David Puertolas, Vasyagina, Ekaterina, Zhu, Kevin, Ferrando, Javier, Chaudhary, Maheep
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915921198579712
author Merenciano, David Puertolas
Vasyagina, Ekaterina
Zhu, Kevin
Ferrando, Javier
Chaudhary, Maheep
author_facet Merenciano, David Puertolas
Vasyagina, Ekaterina
Zhu, Kevin
Ferrando, Javier
Chaudhary, Maheep
contents LoRA adapters let users fine-tune large language models (LLMs) efficiently. However, LoRA adapters are shared through open repositories like Hugging Face Hub \citep{huggingface_hub_docs}, making them vulnerable to backdoor attacks. Current detection methods require running the model with test input data -- making them impractical for screening thousands of adapters where the trigger for backdoor behavior is unknown. We detect poisoned adapters by analyzing their weight matrices directly, without running the model -- making our method trigger-agnostic. For each attention projection (Q, K, V, O), our method extracts five spectral statistics from the low-rank update $ΔW$, yielding a 20-dimensional signature for each adapter. A logistic regression detector trained on this representation separates benign and poisoned adapters across three model families -- Llama-3.2-3B~\citep{llama3}, Qwen2.5-3B~\citep{qwen25}, and Gemma-2-2B~\citep{gemma2} -- on unseen test adapters drawn from instruction-following, reasoning, question-answering, code, and classification tasks. Across all three architectures, the detector achieves 100\% accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15195
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Weight space Detection of Backdoors in LoRA Adapters
Merenciano, David Puertolas
Vasyagina, Ekaterina
Zhu, Kevin
Ferrando, Javier
Chaudhary, Maheep
Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
LoRA adapters let users fine-tune large language models (LLMs) efficiently. However, LoRA adapters are shared through open repositories like Hugging Face Hub \citep{huggingface_hub_docs}, making them vulnerable to backdoor attacks. Current detection methods require running the model with test input data -- making them impractical for screening thousands of adapters where the trigger for backdoor behavior is unknown. We detect poisoned adapters by analyzing their weight matrices directly, without running the model -- making our method trigger-agnostic. For each attention projection (Q, K, V, O), our method extracts five spectral statistics from the low-rank update $ΔW$, yielding a 20-dimensional signature for each adapter. A logistic regression detector trained on this representation separates benign and poisoned adapters across three model families -- Llama-3.2-3B~\citep{llama3}, Qwen2.5-3B~\citep{qwen25}, and Gemma-2-2B~\citep{gemma2} -- on unseen test adapters drawn from instruction-following, reasoning, question-answering, code, and classification tasks. Across all three architectures, the detector achieves 100\% accuracy.
title Weight space Detection of Backdoors in LoRA Adapters
topic Cryptography and Security
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2602.15195