Weight space Detection of Backdoors in LoRA Adapters
Fuente:
arXiv
Saved in:
| Main Authors: | Merenciano, David Puertolas, Vasyagina, Ekaterina, Zhu, Kevin, Ferrando, Javier, Chaudhary, Maheep |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem
by: Liu, Hongyi, et al.
Published: (2024)
by: Liu, Hongyi, et al.
Published: (2024)
LoBAM: LoRA-Based Backdoor Attack on Model Merging
by: Yin, Ming, et al.
Published: (2024)
by: Yin, Ming, et al.
Published: (2024)
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
by: Lelle, Travis
Published: (2026)
by: Lelle, Travis
Published: (2026)
Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA Models
by: Chen, Linzhi, et al.
Published: (2025)
by: Chen, Linzhi, et al.
Published: (2025)
LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models
by: Ran, Delong, et al.
Published: (2025)
by: Ran, Delong, et al.
Published: (2025)
LoRA as Oracle
by: Arazzi, Marco, et al.
Published: (2026)
by: Arazzi, Marco, et al.
Published: (2026)
MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs
by: Kan, Chun Yan Ryan, et al.
Published: (2026)
by: Kan, Chun Yan Ryan, et al.
Published: (2026)
ZKLoRA: Efficient Zero-Knowledge Proofs for LoRA Verification
by: Roy, Bidhan, et al.
Published: (2025)
by: Roy, Bidhan, et al.
Published: (2025)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
by: Batra, Shourya, et al.
Published: (2025)
by: Batra, Shourya, et al.
Published: (2025)
DMFI: A Dual-Modality Log Analysis Framework for Insider Threat Detection with LoRA-Tuned Language Models
by: Kong, Kaichuan, et al.
Published: (2025)
by: Kong, Kaichuan, et al.
Published: (2025)
Universal Jailbreak Backdoors from Poisoned Human Feedback
by: Rando, Javier, et al.
Published: (2023)
by: Rando, Javier, et al.
Published: (2023)
Promoting Data and Model Privacy in Federated Learning through Quantized LoRA
by: Zhu, JianHao, et al.
Published: (2024)
by: Zhu, JianHao, et al.
Published: (2024)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
by: Xu, Jiashu, et al.
Published: (2023)
by: Xu, Jiashu, et al.
Published: (2023)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
by: Rando, Javier, et al.
Published: (2024)
by: Rando, Javier, et al.
Published: (2024)
A Study of Backdoors in Instruction Fine-tuned Language Models
by: Raghuram, Jayaram, et al.
Published: (2024)
by: Raghuram, Jayaram, et al.
Published: (2024)
AC-LoRA: (Almost) Training-Free Access Control-Aware Multi-Modal LLMs
by: Lazier, Lara Magdalena, et al.
Published: (2025)
by: Lazier, Lara Magdalena, et al.
Published: (2025)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Leaner Training, Lower Leakage: Revisiting Memorization in LLM Fine-Tuning with LoRA
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
by: Wang, Guanchu, et al.
Published: (2024)
by: Wang, Guanchu, et al.
Published: (2024)
AuthenLoRA: Entangling Stylization with Imperceptible Watermarks for Copyright-Secure LoRA Adapters
by: Shi, Fangming, et al.
Published: (2025)
by: Shi, Fangming, et al.
Published: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
by: Liu, Qin, et al.
Published: (2023)
by: Liu, Qin, et al.
Published: (2023)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
by: Hu, Man, et al.
Published: (2025)
by: Hu, Man, et al.
Published: (2025)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
by: Tong, Terry, et al.
Published: (2024)
by: Tong, Terry, et al.
Published: (2024)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Fast and Lightweight Backdoor Detection via Head Random Probing
by: Yu, Yinbo, et al.
Published: (2026)
by: Yu, Yinbo, et al.
Published: (2026)
Obliviate: Neutralizing Task-agnostic Backdoors within the Parameter-efficient Fine-tuning Paradigm
by: Kim, Jaehan, et al.
Published: (2024)
by: Kim, Jaehan, et al.
Published: (2024)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
Detecting and Eliminating Neural Network Backdoors Through Active Paths with Application to Intrusion Detection
by: Høyheim, Eirik, et al.
Published: (2026)
by: Høyheim, Eirik, et al.
Published: (2026)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
by: Vega, Jason, et al.
Published: (2023)
by: Vega, Jason, et al.
Published: (2023)
LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters
by: Ahn, Beomjin, et al.
Published: (2026)
by: Ahn, Beomjin, et al.
Published: (2026)
Key-Conditioned Orthonormal Transform Gating (K-OTG): Multi-Key Access Control with Hidden-State Scrambling for LoRA-Tuned Models
by: Khan, Muhammad Haris
Published: (2025)
by: Khan, Muhammad Haris
Published: (2025)
How to Backdoor the Knowledge Distillation
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
LoRA-based Parameter-Efficient LLMs for Continuous Learning in Edge-based Malware Detection
by: Rondanini, Christian, et al.
Published: (2026)
by: Rondanini, Christian, et al.
Published: (2026)
PBP: Post-training Backdoor Purification for Malware Classifiers
by: Nguyen, Dung Thuy, et al.
Published: (2024)
by: Nguyen, Dung Thuy, et al.
Published: (2024)
On the (In)feasibility of ML Backdoor Detection as an Hypothesis Testing Problem
by: Pichler, Georg, et al.
Published: (2024)
by: Pichler, Georg, et al.
Published: (2024)
PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Backdoor Detection through Replicated Execution of Outsourced Training
by: Jia, Hengrui, et al.
Published: (2025)
by: Jia, Hengrui, et al.
Published: (2025)
Exploring Backdoor Vulnerabilities of Chat Models
by: Hao, Yunzhuo, et al.
Published: (2024)
by: Hao, Yunzhuo, et al.
Published: (2024)
Similar Items
-
LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem
by: Liu, Hongyi, et al.
Published: (2024) -
LoBAM: LoRA-Based Backdoor Attack on Model Merging
by: Yin, Ming, et al.
Published: (2024) -
Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
by: Lelle, Travis
Published: (2026) -
Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA Models
by: Chen, Linzhi, et al.
Published: (2025) -
LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models
by: Ran, Delong, et al.
Published: (2025)