Mitigating the Backdoor Effect for Multi-Task Model Merging via Safety-Aware Subspace
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Jinluan, Tang, Anke, Zhu, Didi, Chen, Zhengyu, Shen, Li, Wu, Fei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BadMerging: Backdoor Attacks Against Model Merging
by: Zhang, Jinghuai, et al.
Published: (2024)
by: Zhang, Jinghuai, et al.
Published: (2024)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
by: Min, Rui, et al.
Published: (2024)
by: Min, Rui, et al.
Published: (2024)
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023)
by: Huang, Hai, et al.
Published: (2023)
DMGNN: Detecting and Mitigating Backdoor Attacks in Graph Neural Networks
by: Sui, Hao, et al.
Published: (2024)
by: Sui, Hao, et al.
Published: (2024)
BackdoorBench: A Comprehensive Benchmark and Analysis of Backdoor Learning
by: Wu, Baoyuan, et al.
Published: (2024)
by: Wu, Baoyuan, et al.
Published: (2024)
Diff-Cleanse: Identifying and Mitigating Backdoor Attacks in Diffusion Models
by: Hao, Jiang, et al.
Published: (2024)
by: Hao, Jiang, et al.
Published: (2024)
Fusing Pruned and Backdoored Models: Optimal Transport-based Data-free Backdoor Mitigation
by: Lin, Weilin, et al.
Published: (2024)
by: Lin, Weilin, et al.
Published: (2024)
LoBAM: LoRA-Based Backdoor Attack on Model Merging
by: Yin, Ming, et al.
Published: (2024)
by: Yin, Ming, et al.
Published: (2024)
Self-Purification Mitigates Backdoors in Multimodal Diffusion Language Models
by: Wan, Guangnian, et al.
Published: (2026)
by: Wan, Guangnian, et al.
Published: (2026)
Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
by: Yuan, Zenghui, et al.
Published: (2025)
by: Yuan, Zenghui, et al.
Published: (2025)
BADTV: Unveiling Backdoor Threats in Third-Party Task Vectors
by: Hsu, Chia-Yi, et al.
Published: (2025)
by: Hsu, Chia-Yi, et al.
Published: (2025)
Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data
by: Zhang, Jiale, et al.
Published: (2025)
by: Zhang, Jiale, et al.
Published: (2025)
"No Matter What You Do": Purifying GNN Models via Backdoor Unlearning
by: Zhang, Jiale, et al.
Published: (2024)
by: Zhang, Jiale, et al.
Published: (2024)
Differentially Private Subspace Fine-Tuning for Large Language Models
by: Zheng, Lele, et al.
Published: (2026)
by: Zheng, Lele, et al.
Published: (2026)
Gradient Shaping: Enhancing Backdoor Attack Against Reverse Engineering
by: Zhu, Rui, et al.
Published: (2023)
by: Zhu, Rui, et al.
Published: (2023)
Awakening the Hydra: Stabilizing Multi-Concept Backdoor Injection in Text-to-Image Diffusion Models
by: Wang, Kai, et al.
Published: (2026)
by: Wang, Kai, et al.
Published: (2026)
CENSOR: Defense Against Gradient Inversion via Orthogonal Subspace Bayesian Sampling
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
Structure-Aware Distributed Backdoor Attacks in Federated Learning
by: Jian, Wang, et al.
Published: (2026)
by: Jian, Wang, et al.
Published: (2026)
Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution
by: Wu, Yixin, et al.
Published: (2024)
by: Wu, Yixin, et al.
Published: (2024)
How to Backdoor the Knowledge Distillation
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
Countering Backdoor Attacks in Image Recognition: A Survey and Evaluation of Mitigation Strategies
by: Dunnett, Kealan, et al.
Published: (2024)
by: Dunnett, Kealan, et al.
Published: (2024)
BadRSSD: Backdoor Attacks on Regularized Self-Supervised Diffusion Models
by: Wang, Jiayao, et al.
Published: (2026)
by: Wang, Jiayao, et al.
Published: (2026)
Backdooring Masked Diffusion Language Models
by: Cao, Daniel Yiming, et al.
Published: (2026)
by: Cao, Daniel Yiming, et al.
Published: (2026)
Unelicitable Backdoors in Language Models via Cryptographic Transformer Circuits
by: Draguns, Andis, et al.
Published: (2024)
by: Draguns, Andis, et al.
Published: (2024)
DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
by: Qiu, Han, et al.
Published: (2020)
by: Qiu, Han, et al.
Published: (2020)
Instruction Backdoor Attacks Against Customized LLMs
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
by: Pawlak, Stanisław, et al.
Published: (2025)
by: Pawlak, Stanisław, et al.
Published: (2025)
Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNs
by: Yang, Han, et al.
Published: (2025)
by: Yang, Han, et al.
Published: (2025)
Rethinking Pruning for Backdoor Mitigation: An Optimization Perspective
by: Li, Nan, et al.
Published: (2024)
by: Li, Nan, et al.
Published: (2024)
On the Robustness of Graph Reduction Against GNN Backdoor
by: Zhu, Yuxuan, et al.
Published: (2024)
by: Zhu, Yuxuan, et al.
Published: (2024)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
by: Liu, Qin, et al.
Published: (2024)
by: Liu, Qin, et al.
Published: (2024)
Backdoor Mitigation in Deep Neural Networks via Strategic Retraining
by: Dhonthi, Akshay, et al.
Published: (2022)
by: Dhonthi, Akshay, et al.
Published: (2022)
Mutual Information Guided Backdoor Mitigation for Pre-trained Encoders
by: Han, Tingxu, et al.
Published: (2024)
by: Han, Tingxu, et al.
Published: (2024)
Graph-Aware Stealthy Poison-Text Backdoors for Text-Attributed Graphs
by: Luo, Qi, et al.
Published: (2026)
by: Luo, Qi, et al.
Published: (2026)
Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift
by: An, Shengwei, et al.
Published: (2023)
by: An, Shengwei, et al.
Published: (2023)
ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
by: Ren, Zhiyao, et al.
Published: (2025)
by: Ren, Zhiyao, et al.
Published: (2025)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
by: Xu, Jiashu, et al.
Published: (2023)
by: Xu, Jiashu, et al.
Published: (2023)
PiDAn: A Coherence Optimization Approach for Backdoor Attack Detection and Mitigation in Deep Neural Networks
by: Wang, Yue, et al.
Published: (2022)
by: Wang, Yue, et al.
Published: (2022)
FedBAP: Backdoor Defense via Benign Adversarial Perturbation in Federated Learning
by: Yan, Xinhai, et al.
Published: (2025)
by: Yan, Xinhai, et al.
Published: (2025)
Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs
by: Lu, Yiyang, et al.
Published: (2026)
by: Lu, Yiyang, et al.
Published: (2026)
Similar Items
-
BadMerging: Backdoor Attacks Against Model Merging
by: Zhang, Jinghuai, et al.
Published: (2024) -
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
by: Min, Rui, et al.
Published: (2024) -
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023) -
DMGNN: Detecting and Mitigating Backdoor Attacks in Graph Neural Networks
by: Sui, Hao, et al.
Published: (2024) -
BackdoorBench: A Comprehensive Benchmark and Analysis of Backdoor Learning
by: Wu, Baoyuan, et al.
Published: (2024)