Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Beniwal, Himanshu, Panda, Sailesh, Srivibhav, Birudugadda, Singh, Mayank |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-lingual Editing in Multilingual Language Models
by: Beniwal, Himanshu, et al.
Published: (2024)
by: Beniwal, Himanshu, et al.
Published: (2024)
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
by: Beniwal, Himanshu, et al.
Published: (2025)
by: Beniwal, Himanshu, et al.
Published: (2025)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
by: Yadav, Ankit, et al.
Published: (2024)
by: Yadav, Ankit, et al.
Published: (2024)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
by: Beniwal, Himanshu, et al.
Published: (2026)
by: Beniwal, Himanshu, et al.
Published: (2026)
COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
by: Sheth, Rajvee, et al.
Published: (2024)
by: Sheth, Rajvee, et al.
Published: (2024)
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
by: Sheth, Rajvee, et al.
Published: (2025)
by: Sheth, Rajvee, et al.
Published: (2025)
DEPART: DEcomposing PARiTy across Multilingual LLMs
by: Uppadhyay, Manan, et al.
Published: (2026)
by: Uppadhyay, Manan, et al.
Published: (2026)
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
by: Feng, Yunhao, et al.
Published: (2026)
by: Feng, Yunhao, et al.
Published: (2026)
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
by: Beniwal, Himanshu, et al.
Published: (2025)
by: Beniwal, Himanshu, et al.
Published: (2025)
Debiasing Multilingual LLMs in Cross-lingual Latent Space
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
by: Han, HyoJung, et al.
Published: (2025)
by: Han, HyoJung, et al.
Published: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
by: Wei, Jiali, et al.
Published: (2026)
by: Wei, Jiali, et al.
Published: (2026)
Cross-lingual Transfer of Reward Models in Multilingual Alignment
by: Hong, Jiwoo, et al.
Published: (2024)
by: Hong, Jiwoo, et al.
Published: (2024)
When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models
by: Panda, Sailesh, et al.
Published: (2026)
by: Panda, Sailesh, et al.
Published: (2026)
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Remember This Event That Year? Assessing Temporal Information and Reasoning in Large Language Models
by: Beniwal, Himanshu, et al.
Published: (2024)
by: Beniwal, Himanshu, et al.
Published: (2024)
MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation
by: Liu, Yile, et al.
Published: (2025)
by: Liu, Yile, et al.
Published: (2025)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
One Instruction Does Not Fit All: How Well Do Embeddings Align Personas and Instructions in Low-Resource Indian Languages?
by: Shah, Arya, et al.
Published: (2026)
by: Shah, Arya, et al.
Published: (2026)
Claim-Guided Textual Backdoor Attack for Practical Applications
by: Song, Minkyoo, et al.
Published: (2024)
by: Song, Minkyoo, et al.
Published: (2024)
BadCLM: Backdoor Attack in Clinical Language Models for Electronic Health Records
by: Lyu, Weimin, et al.
Published: (2024)
by: Lyu, Weimin, et al.
Published: (2024)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
by: Du, Wei, et al.
Published: (2023)
by: Du, Wei, et al.
Published: (2023)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Cross-lingual QA: A Key to Unlocking In-context Cross-lingual Performance
by: Kim, Sunkyoung, et al.
Published: (2023)
by: Kim, Sunkyoung, et al.
Published: (2023)
VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
by: Ye, Ziang, et al.
Published: (2025)
by: Ye, Ziang, et al.
Published: (2025)
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
by: Shimabucoro, Luisa, et al.
Published: (2025)
by: Shimabucoro, Luisa, et al.
Published: (2025)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
by: Jiang, Peihai, et al.
Published: (2025)
by: Jiang, Peihai, et al.
Published: (2025)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
by: Ji, Wence, et al.
Published: (2025)
by: Ji, Wence, et al.
Published: (2025)
xCoT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought Reasoning
by: Chai, Linzheng, et al.
Published: (2024)
by: Chai, Linzheng, et al.
Published: (2024)
When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models
by: Liu, Ziyu, et al.
Published: (2026)
by: Liu, Ziyu, et al.
Published: (2026)
Multilingual Information Retrieval with a Monolingual Knowledge Base
by: Zhuang, Yingying, et al.
Published: (2025)
by: Zhuang, Yingying, et al.
Published: (2025)
MPN: Leveraging Multilingual Patch Neuron for Cross-lingual Model Editing
by: Si, Nianwen, et al.
Published: (2024)
by: Si, Nianwen, et al.
Published: (2024)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
by: Ifergan, Maxim, et al.
Published: (2024)
by: Ifergan, Maxim, et al.
Published: (2024)
What Drives Cross-lingual Ranking? Retrieval Approaches with Multilingual Language Models
by: Goworek, Roksana, et al.
Published: (2025)
by: Goworek, Roksana, et al.
Published: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
by: Cheng, Pengzhou, et al.
Published: (2024)
by: Cheng, Pengzhou, et al.
Published: (2024)
Similar Items
-
Cross-lingual Editing in Multilingual Language Models
by: Beniwal, Himanshu, et al.
Published: (2024) -
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
by: Beniwal, Himanshu, et al.
Published: (2025) -
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
by: Yadav, Ankit, et al.
Published: (2024) -
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
by: Beniwal, Himanshu, et al.
Published: (2026) -
COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
by: Sheth, Rajvee, et al.
Published: (2024)