Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Jiashu, Ma, Mingyu Derek, Wang, Fei, Xiao, Chaowei, Chen, Muhao |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Instructional Fingerprinting of Large Language Models
by: Xu, Jiashu, et al.
Published: (2024)
by: Xu, Jiashu, et al.
Published: (2024)
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
by: Liu, Qin, et al.
Published: (2024)
by: Liu, Qin, et al.
Published: (2024)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
by: Liu, Qin, et al.
Published: (2023)
by: Liu, Qin, et al.
Published: (2023)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
by: Tong, Terry, et al.
Published: (2024)
by: Tong, Terry, et al.
Published: (2024)
A Study of Backdoors in Instruction Fine-tuned Language Models
by: Raghuram, Jayaram, et al.
Published: (2024)
by: Raghuram, Jayaram, et al.
Published: (2024)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025)
by: Tong, Terry, et al.
Published: (2025)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
by: Yan, Jun, et al.
Published: (2023)
by: Yan, Jun, et al.
Published: (2023)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Exploring Backdoor Vulnerabilities of Chat Models
by: Hao, Yunzhuo, et al.
Published: (2024)
by: Hao, Yunzhuo, et al.
Published: (2024)
How to Backdoor the Knowledge Distillation
by: Wu, Chen, et al.
Published: (2025)
by: Wu, Chen, et al.
Published: (2025)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
by: Foroughi, Mohammad Hadi, et al.
Published: (2026)
Universal Jailbreak Backdoors from Poisoned Human Feedback
by: Rando, Javier, et al.
Published: (2023)
by: Rando, Javier, et al.
Published: (2023)
Weight space Detection of Backdoors in LoRA Adapters
by: Merenciano, David Puertolas, et al.
Published: (2026)
by: Merenciano, David Puertolas, et al.
Published: (2026)
Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Model Watermarking
by: Kong, Cong, et al.
Published: (2024)
by: Kong, Cong, et al.
Published: (2024)
Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models
by: Shin, Jeongjin, et al.
Published: (2024)
by: Shin, Jeongjin, et al.
Published: (2024)
Backdoor defense, learnability and obfuscation
by: Christiano, Paul, et al.
Published: (2024)
by: Christiano, Paul, et al.
Published: (2024)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
by: Betley, Jan, et al.
Published: (2025)
by: Betley, Jan, et al.
Published: (2025)
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
by: Rando, Javier, et al.
Published: (2024)
by: Rando, Javier, et al.
Published: (2024)
Invisible Backdoor Attack Through Singular Value Decomposition
by: Chen, Wenmin, et al.
Published: (2024)
by: Chen, Wenmin, et al.
Published: (2024)
Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses
by: Pawlak, Stanisław, et al.
Published: (2025)
by: Pawlak, Stanisław, et al.
Published: (2025)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
by: Wu, Tong, et al.
Published: (2024)
by: Wu, Tong, et al.
Published: (2024)
Heterogeneous Graph Backdoor Attack
by: Chen, Jiawei, et al.
Published: (2025)
by: Chen, Jiawei, et al.
Published: (2025)
Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution Shift
by: An, Shengwei, et al.
Published: (2023)
by: An, Shengwei, et al.
Published: (2023)
Backdoor Graph Condensation
by: Wu, Jiahao, et al.
Published: (2024)
by: Wu, Jiahao, et al.
Published: (2024)
Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency
by: Pal, Soumyadeep, et al.
Published: (2024)
by: Pal, Soumyadeep, et al.
Published: (2024)
How to Craft Backdoors with Unlabeled Data Alone?
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
by: Cao, Yuanpu, et al.
Published: (2023)
by: Cao, Yuanpu, et al.
Published: (2023)
Obliviate: Neutralizing Task-agnostic Backdoors within the Parameter-efficient Fine-tuning Paradigm
by: Kim, Jaehan, et al.
Published: (2024)
by: Kim, Jaehan, et al.
Published: (2024)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
by: Xue, Eric, et al.
Published: (2025)
by: Xue, Eric, et al.
Published: (2025)
Rounding-Guided Backdoor Injection in Deep Learning Model Quantization
by: Chen, Xiangxiang, et al.
Published: (2025)
by: Chen, Xiangxiang, et al.
Published: (2025)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
BACKTIME: Backdoor Attacks on Multivariate Time Series Forecasting
by: Lin, Xiao, et al.
Published: (2024)
by: Lin, Xiao, et al.
Published: (2024)
Injecting Universal Jailbreak Backdoors into LLMs in Minutes
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
by: Liu, Fuqiang, et al.
Published: (2024)
by: Liu, Fuqiang, et al.
Published: (2024)
Graph Neural Backdoor: Fundamentals, Methodologies, Applications, and Future Directions
by: Yang, Xiao, et al.
Published: (2024)
by: Yang, Xiao, et al.
Published: (2024)
Composite Backdoor Attacks Against Large Language Models
by: Huang, Hai, et al.
Published: (2023)
by: Huang, Hai, et al.
Published: (2023)
BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets
by: Gong, Chen, et al.
Published: (2022)
by: Gong, Chen, et al.
Published: (2022)
Backdooring Bias ($B^2$) into Stable Diffusion Models
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
Similar Items
-
Instructional Fingerprinting of Large Language Models
by: Xu, Jiashu, et al.
Published: (2024) -
Mitigating Backdoor Threats to Large Language Models: Advancement and Challenges
by: Liu, Qin, et al.
Published: (2024) -
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
by: Liu, Qin, et al.
Published: (2023) -
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
by: Tong, Terry, et al.
Published: (2024) -
A Study of Backdoors in Instruction Fine-tuned Language Models
by: Raghuram, Jayaram, et al.
Published: (2024)