Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Cui, Jing, Han, Yufei, Jiao, Jianbin, Zhang, Junge |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Recent Advances in Attack and Defense Approaches of Large Language Models
di: Cui, Jing, et al.
Pubblicazione: (2024)
di: Cui, Jing, et al.
Pubblicazione: (2024)
Towards Effective, Stealthy, and Persistent Backdoor Attacks Targeting Graph Foundation Models
di: Luo, Jiayi, et al.
Pubblicazione: (2025)
di: Luo, Jiayi, et al.
Pubblicazione: (2025)
Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors
di: Li, Zi, et al.
Pubblicazione: (2026)
di: Li, Zi, et al.
Pubblicazione: (2026)
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
di: Guo, Weiyang, et al.
Pubblicazione: (2026)
di: Guo, Weiyang, et al.
Pubblicazione: (2026)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
Scam Shield: Multi-Model Voting and Fine-Tuned LLMs Against Adversarial Attacks
di: Chang, Chen-Wei, et al.
Pubblicazione: (2025)
di: Chang, Chen-Wei, et al.
Pubblicazione: (2025)
Stealthy Backdoor Attack to Real-world Models in Android Apps
di: Wei, Jiali, et al.
Pubblicazione: (2025)
di: Wei, Jiali, et al.
Pubblicazione: (2025)
Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models
di: Xue, Xiaoyu, et al.
Pubblicazione: (2025)
di: Xue, Xiaoyu, et al.
Pubblicazione: (2025)
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
di: Zhang, Rui, et al.
Pubblicazione: (2025)
di: Zhang, Rui, et al.
Pubblicazione: (2025)
Invisible Textual Backdoor Attacks based on Dual-Trigger
di: Hou, Yang, et al.
Pubblicazione: (2024)
di: Hou, Yang, et al.
Pubblicazione: (2024)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
di: Li, Yige, et al.
Pubblicazione: (2025)
di: Li, Yige, et al.
Pubblicazione: (2025)
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
di: Zhou, Yihe, et al.
Pubblicazione: (2025)
di: Zhou, Yihe, et al.
Pubblicazione: (2025)
When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
di: Ge, Huaizhi, et al.
Pubblicazione: (2024)
di: Ge, Huaizhi, et al.
Pubblicazione: (2024)
FTSmartAudit: A Knowledge Distillation-Enhanced Framework for Automated Smart Contract Auditing Using Fine-Tuned LLMs
di: Wei, Zhiyuan, et al.
Pubblicazione: (2024)
di: Wei, Zhiyuan, et al.
Pubblicazione: (2024)
Enhancing All-to-X Backdoor Attacks with Optimized Target Class Mapping
di: Wang, Lei, et al.
Pubblicazione: (2025)
di: Wang, Lei, et al.
Pubblicazione: (2025)
Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs
di: Wang, Yifei, et al.
Pubblicazione: (2026)
di: Wang, Yifei, et al.
Pubblicazione: (2026)
Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes
di: Kumar, Divyanshu, et al.
Pubblicazione: (2024)
di: Kumar, Divyanshu, et al.
Pubblicazione: (2024)
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
di: Li, Yige, et al.
Pubblicazione: (2026)
di: Li, Yige, et al.
Pubblicazione: (2026)
Flashy Backdoor: Real-world Environment Backdoor Attack on SNNs with DVS Cameras
di: Riaño, Roberto, et al.
Pubblicazione: (2024)
di: Riaño, Roberto, et al.
Pubblicazione: (2024)
BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting
di: Qiu, Huming, et al.
Pubblicazione: (2023)
di: Qiu, Huming, et al.
Pubblicazione: (2023)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
di: Wei, Jiali, et al.
Pubblicazione: (2026)
di: Wei, Jiali, et al.
Pubblicazione: (2026)
Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning
di: Song, Baogang, et al.
Pubblicazione: (2025)
di: Song, Baogang, et al.
Pubblicazione: (2025)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
di: Zheng, Jingyi, et al.
Pubblicazione: (2024)
di: Zheng, Jingyi, et al.
Pubblicazione: (2024)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
di: Poppi, Samuele, et al.
Pubblicazione: (2024)
di: Poppi, Samuele, et al.
Pubblicazione: (2024)
CUBA: Controlled Untargeted Backdoor Attack against Deep Neural Networks
di: Wu, Yinghao, et al.
Pubblicazione: (2025)
di: Wu, Yinghao, et al.
Pubblicazione: (2025)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
di: Li, Xi, et al.
Pubblicazione: (2024)
di: Li, Xi, et al.
Pubblicazione: (2024)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
di: Zhao, Shuai, et al.
Pubblicazione: (2024)
SSCL-BW: Sample-Specific Clean-Label Backdoor Watermarking for Dataset Ownership Verification
di: Wang, Yingjia, et al.
Pubblicazione: (2025)
di: Wang, Yingjia, et al.
Pubblicazione: (2025)
DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
di: Popovic, Dorde, et al.
Pubblicazione: (2025)
di: Popovic, Dorde, et al.
Pubblicazione: (2025)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
di: Liu, Xuxu, et al.
Pubblicazione: (2025)
di: Liu, Xuxu, et al.
Pubblicazione: (2025)
From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
di: Shen, Guangyu, et al.
Pubblicazione: (2025)
di: Shen, Guangyu, et al.
Pubblicazione: (2025)
Concept-Guided Backdoor Attack on Vision Language Models
di: Shen, Haoyu, et al.
Pubblicazione: (2025)
di: Shen, Haoyu, et al.
Pubblicazione: (2025)
Revisiting Training-Inference Trigger Intensity in Backdoor Attacks
di: Lin, Chenhao, et al.
Pubblicazione: (2025)
di: Lin, Chenhao, et al.
Pubblicazione: (2025)
Exploring Backdoor Attack and Defense for LLM-empowered Recommendations
di: Ning, Liangbo, et al.
Pubblicazione: (2025)
di: Ning, Liangbo, et al.
Pubblicazione: (2025)
Differentiation-Based Extraction of Proprietary Data from Fine-Tuned LLMs
di: Li, Zongjie, et al.
Pubblicazione: (2025)
di: Li, Zongjie, et al.
Pubblicazione: (2025)
Clean-Label Physical Backdoor Attacks with Data Distillation
di: Dao, Thinh, et al.
Pubblicazione: (2024)
di: Dao, Thinh, et al.
Pubblicazione: (2024)
Impart: An Imperceptible and Effective Label-Specific Backdoor Attack
di: Zhao, Jingke, et al.
Pubblicazione: (2024)
di: Zhao, Jingke, et al.
Pubblicazione: (2024)
PAD-FT: A Lightweight Defense for Backdoor Attacks via Data Purification and Fine-Tuning
di: Xu, Yukai, et al.
Pubblicazione: (2024)
di: Xu, Yukai, et al.
Pubblicazione: (2024)
PEFT-as-an-Attack! Jailbreaking Language Models during Federated Parameter-Efficient Fine-Tuning
di: Li, Shenghui, et al.
Pubblicazione: (2024)
di: Li, Shenghui, et al.
Pubblicazione: (2024)
Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models
di: Ding, Zikang, et al.
Pubblicazione: (2026)
di: Ding, Zikang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Recent Advances in Attack and Defense Approaches of Large Language Models
di: Cui, Jing, et al.
Pubblicazione: (2024) -
Towards Effective, Stealthy, and Persistent Backdoor Attacks Targeting Graph Foundation Models
di: Luo, Jiayi, et al.
Pubblicazione: (2025) -
Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors
di: Li, Zi, et al.
Pubblicazione: (2026) -
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
di: Guo, Weiyang, et al.
Pubblicazione: (2026) -
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
di: Zhao, Shuai, et al.
Pubblicazione: (2024)