TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | He, Xuanli, Wang, Jun, Xu, Qiongkai, Minervini, Pasquale, Stenetorp, Pontus, Rubinstein, Benjamin I. P., Cohn, Trevor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
by: He, Xuanli, et al.
Published: (2024)
by: He, Xuanli, et al.
Published: (2024)
Defending against Backdoor Attacks via Module Switching
by: Li, Weijun, et al.
Published: (2025)
by: Li, Weijun, et al.
Published: (2025)
Cut the Deadwood Out: Backdoor Purification via Guided Module Substitution
by: Tong, Yao, et al.
Published: (2024)
by: Tong, Yao, et al.
Published: (2024)
Backdoor Attack on Multilingual Machine Translation
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)
by: Zheng, Jingyi, et al.
Published: (2024)
Generative Models are Self-Watermarked: Declaring Model Authentication through Re-Generation
by: Desu, Aditya, et al.
Published: (2024)
by: Desu, Aditya, et al.
Published: (2024)
Instruction Backdoor Attacks Against Customized LLMs
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
Attacks on Third-Party APIs of Large Language Models
by: Zhao, Wanru, et al.
Published: (2024)
by: Zhao, Wanru, et al.
Published: (2024)
Double Backdoored: Converting Code Large Language Model Backdoors to Traditional Malware via Adversarial Instruction Tuning Attacks
by: Hossen, Md Imran, et al.
Published: (2024)
by: Hossen, Md Imran, et al.
Published: (2024)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
by: Cui, Jing, et al.
Published: (2025)
by: Cui, Jing, et al.
Published: (2025)
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
by: Sun, Zhen, et al.
Published: (2024)
by: Sun, Zhen, et al.
Published: (2024)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
by: Wen, Rui, et al.
Published: (2026)
by: Wen, Rui, et al.
Published: (2026)
ProtegoFed: Backdoor-Free Federated Instruction Tuning with Interspersed Poisoned Data
by: Zhao, Haodong, et al.
Published: (2026)
by: Zhao, Haodong, et al.
Published: (2026)
WARDEN: Multi-Directional Backdoor Watermarks for Embedding-as-a-Service Copyright Protection
by: Shetty, Anudeex, et al.
Published: (2024)
by: Shetty, Anudeex, et al.
Published: (2024)
Getting a-Round Guarantees: Floating-Point Attacks on Certified Robustness
by: Jin, Jiankai, et al.
Published: (2022)
by: Jin, Jiankai, et al.
Published: (2022)
The Invitation Trap: Proactive Availability Backdoor in LLMs via Conversational Induction
by: Wang, He, et al.
Published: (2026)
by: Wang, He, et al.
Published: (2026)
Revisiting Backdoor Threat in Federated Instruction Tuning from a Signal Aggregation Perspective
by: Zhao, Haodong, et al.
Published: (2026)
by: Zhao, Haodong, et al.
Published: (2026)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
by: Fu, Yu, et al.
Published: (2024)
by: Fu, Yu, et al.
Published: (2024)
Krait: A Backdoor Attack Against Graph Prompt Tuning
by: Song, Ying, et al.
Published: (2024)
by: Song, Ying, et al.
Published: (2024)
Are We There Yet? Timing and Floating-Point Attacks on Differential Privacy Systems
by: Jin, Jiankai, et al.
Published: (2021)
by: Jin, Jiankai, et al.
Published: (2021)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
by: Xu, Jiashu, et al.
Published: (2023)
by: Xu, Jiashu, et al.
Published: (2023)
Robust Anti-Backdoor Instruction Tuning in LVLMs
by: Xun, Yuan, et al.
Published: (2025)
by: Xun, Yuan, et al.
Published: (2025)
Stealthy Backdoor Attack via Confidence-driven Sampling
by: He, Pengfei, et al.
Published: (2023)
by: He, Pengfei, et al.
Published: (2023)
Cross-Lingual Summarization as a Black-Box Watermark Removal Attack
by: Ganesan, Gokul
Published: (2025)
by: Ganesan, Gokul
Published: (2025)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
iBA: Backdoor Attack on 3D Point Cloud via Reconstructing Itself
by: Bian, Yuhao, et al.
Published: (2024)
by: Bian, Yuhao, et al.
Published: (2024)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Combinational Backdoor Attack against Customized Text-to-Image Models
by: Jiang, Wenbo, et al.
Published: (2024)
by: Jiang, Wenbo, et al.
Published: (2024)
Meme Trojan: Backdoor Attacks Against Hateful Meme Detection via Cross-Modal Triggers
by: Wang, Ruofei, et al.
Published: (2024)
by: Wang, Ruofei, et al.
Published: (2024)
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
by: Yan, Jun, et al.
Published: (2023)
by: Yan, Jun, et al.
Published: (2023)
PointBA: Towards Backdoor Attacks in 3D Point Cloud
by: Li, Xinke, et al.
Published: (2021)
by: Li, Xinke, et al.
Published: (2021)
Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks
by: Li, Yige, et al.
Published: (2024)
by: Li, Yige, et al.
Published: (2024)
Cross-Paradigm Graph Backdoor Attacks with Promptable Subgraph Triggers
by: Liu, Dongyi, et al.
Published: (2025)
by: Liu, Dongyi, et al.
Published: (2025)
Let's Focus: Focused Backdoor Attack against Federated Transfer Learning
by: Arazzi, Marco, et al.
Published: (2024)
by: Arazzi, Marco, et al.
Published: (2024)
CIS-BA: Continuous Interaction Space Based Backdoor Attack for Object Detection in the Real-World
by: Zhao, Shuxin, et al.
Published: (2025)
by: Zhao, Shuxin, et al.
Published: (2025)
Target Attack Backdoor Malware Analysis and Attribution
by: Lai, Anthony Cheuk Tung, et al.
Published: (2025)
by: Lai, Anthony Cheuk Tung, et al.
Published: (2025)
TrojanEdit: Multimodal Backdoor Attack Against Image Editing Model
by: Guo, Ji, et al.
Published: (2024)
by: Guo, Ji, et al.
Published: (2024)
Segment-Level Coherence for Robust Harmful Intent Probing in LLMs
by: He, Xuanli, et al.
Published: (2026)
by: He, Xuanli, et al.
Published: (2026)
Similar Items
-
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
by: He, Xuanli, et al.
Published: (2024) -
Defending against Backdoor Attacks via Module Switching
by: Li, Weijun, et al.
Published: (2025) -
Cut the Deadwood Out: Backdoor Purification via Guided Module Substitution
by: Tong, Yao, et al.
Published: (2024) -
Backdoor Attack on Multilingual Machine Translation
by: Wang, Jun, et al.
Published: (2024) -
CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers
by: Zheng, Jingyi, et al.
Published: (2024)