Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Shuai, Wu, Xiaobao, Nguyen, Cong-Duy, Jia, Yanhao, Jia, Meihuizi, Feng, Yichao, Tuan, Luu Anh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
by: Zhao, Shuai, et al.
Published: (2025)
by: Zhao, Shuai, et al.
Published: (2025)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
by: Hu, Man, et al.
Published: (2025)
by: Hu, Man, et al.
Published: (2025)
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
by: Hu, Man, et al.
Published: (2025)
by: Hu, Man, et al.
Published: (2025)
Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models
by: Wen, Jinming, et al.
Published: (2025)
by: Wen, Jinming, et al.
Published: (2025)
Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning
by: Song, Baogang, et al.
Published: (2025)
by: Song, Baogang, et al.
Published: (2025)
When Forgetting Triggers Backdoors: A Clean Unlearning Attack
by: Arazzi, Marco, et al.
Published: (2025)
by: Arazzi, Marco, et al.
Published: (2025)
Robust Knowledge Distillation in Federated Learning: Counteracting Backdoor Attacks
by: Alharbi, Ebtisaam, et al.
Published: (2025)
by: Alharbi, Ebtisaam, et al.
Published: (2025)
Transferring Backdoors between Large Language Models by Knowledge Distillation
by: Cheng, Pengzhou, et al.
Published: (2024)
by: Cheng, Pengzhou, et al.
Published: (2024)
Stealthy Yet Effective: Distribution-Preserving Backdoor Attacks on Graph Classification
by: Wang, Xiaobao, et al.
Published: (2025)
by: Wang, Xiaobao, et al.
Published: (2025)
Distributed Backdoor Attacks on Federated Graph Learning and Certified Defenses
by: Yang, Yuxin, et al.
Published: (2024)
by: Yang, Yuxin, et al.
Published: (2024)
How to Backdoor the Knowledge Distillation
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
Unlearning Backdoor Attacks through Gradient-Based Model Pruning
by: Dunnett, Kealan, et al.
Published: (2024)
by: Dunnett, Kealan, et al.
Published: (2024)
BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
by: Wang, Shanmin, et al.
Published: (2025)
by: Wang, Shanmin, et al.
Published: (2025)
BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation
by: Wu, Zhengxian, et al.
Published: (2025)
by: Wu, Zhengxian, et al.
Published: (2025)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
by: Shin, Jeongjin, et al.
Published: (2024)
by: Shin, Jeongjin, et al.
Published: (2024)
Clean-Label Physical Backdoor Attacks with Data Distillation
by: Dao, Thinh, et al.
Published: (2024)
by: Dao, Thinh, et al.
Published: (2024)
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
by: Su, Yanghao, et al.
Published: (2025)
by: Su, Yanghao, et al.
Published: (2025)
Instruction Backdoor Attacks Against Customized LLMs
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
Synthesizing Physical Backdoor Datasets: An Automated Framework Leveraging Deep Generative Models
by: Yang, Sze Jue, et al.
Published: (2023)
by: Yang, Sze Jue, et al.
Published: (2023)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)
by: Zheng, Jingyi, et al.
Published: (2024)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
by: Wen, Rui, et al.
Published: (2026)
by: Wen, Rui, et al.
Published: (2026)
Reconstruction Attacks on Machine Unlearning: Simple Models are Vulnerable
by: Bertran, Martin, et al.
Published: (2024)
by: Bertran, Martin, et al.
Published: (2024)
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
by: Sun, Zhen, et al.
Published: (2024)
by: Sun, Zhen, et al.
Published: (2024)
Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
HPE: Hallucinated Positive Entanglement for Backdoor Attacks in Federated Self-Supervised Learning
by: Wang, Jiayao, et al.
Published: (2026)
by: Wang, Jiayao, et al.
Published: (2026)
Unlearning-Enhanced Website Fingerprinting Attack: Against Backdoor Poisoning in Anonymous Networks
by: Yuan, Yali, et al.
Published: (2025)
by: Yuan, Yali, et al.
Published: (2025)
KDMCSE: Knowledge Distillation Multimodal Sentence Embeddings with Adaptive Angular margin Contrastive Learning
by: Nguyen, Cong-Duy, et al.
Published: (2024)
by: Nguyen, Cong-Duy, et al.
Published: (2024)
Does Few-shot Learning Suffer from Backdoor Attacks?
by: Liu, Xinwei, et al.
Published: (2023)
by: Liu, Xinwei, et al.
Published: (2023)
LMEraser: Large Model Unlearning through Adaptive Prompt Tuning
by: Xu, Jie, et al.
Published: (2024)
by: Xu, Jie, et al.
Published: (2024)
Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning
by: Hu, Hongsheng, et al.
Published: (2024)
by: Hu, Hongsheng, et al.
Published: (2024)
DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
by: An, Hyeseon, et al.
Published: (2025)
by: An, Hyeseon, et al.
Published: (2025)
Infighting in the Dark: Multi-Label Backdoor Attack in Federated Learning
by: Li, Ye, et al.
Published: (2024)
by: Li, Ye, et al.
Published: (2024)
On the Credibility of Backdoor Attacks Against Object Detectors in the Physical World
by: Doan, Bao Gia, et al.
Published: (2024)
by: Doan, Bao Gia, et al.
Published: (2024)
BDPFL: Backdoor Defense for Personalized Federated Learning via Explainable Distillation
by: Zhu, Chengcheng, et al.
Published: (2025)
by: Zhu, Chengcheng, et al.
Published: (2025)
Backdoor Attacks and Countermeasures in Natural Language Processing Models: A Comprehensive Security Review
by: Cheng, Pengzhou, et al.
Published: (2023)
by: Cheng, Pengzhou, et al.
Published: (2023)
Similar Items
-
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
by: Zhao, Shuai, et al.
Published: (2024) -
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
by: Zhao, Shuai, et al.
Published: (2024) -
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
by: Zhao, Shuai, et al.
Published: (2024) -
P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs
by: Zhao, Shuai, et al.
Published: (2025) -
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)