Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Beniwal, Himanshu, Panda, Sailesh, Srivibhav, Birudugadda, Singh, Mayank |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Cross-lingual Editing in Multilingual Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2024)
por: Beniwal, Himanshu, et al.
Publicado: (2024)
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
por: Beniwal, Himanshu, et al.
Publicado: (2025)
por: Beniwal, Himanshu, et al.
Publicado: (2025)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
por: Yadav, Ankit, et al.
Publicado: (2024)
por: Yadav, Ankit, et al.
Publicado: (2024)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2026)
por: Beniwal, Himanshu, et al.
Publicado: (2026)
COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
por: Sheth, Rajvee, et al.
Publicado: (2024)
por: Sheth, Rajvee, et al.
Publicado: (2024)
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
por: Sheth, Rajvee, et al.
Publicado: (2025)
por: Sheth, Rajvee, et al.
Publicado: (2025)
DEPART: DEcomposing PARiTy across Multilingual LLMs
por: Uppadhyay, Manan, et al.
Publicado: (2026)
por: Uppadhyay, Manan, et al.
Publicado: (2026)
BackdoorAgent: A Unified Framework for Backdoor Attacks on LLM-based Agents
por: Feng, Yunhao, et al.
Publicado: (2026)
por: Feng, Yunhao, et al.
Publicado: (2026)
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
por: Beniwal, Himanshu, et al.
Publicado: (2025)
por: Beniwal, Himanshu, et al.
Publicado: (2025)
Debiasing Multilingual LLMs in Cross-lingual Latent Space
por: Peng, Qiwei, et al.
Publicado: (2025)
por: Peng, Qiwei, et al.
Publicado: (2025)
How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
por: Gupta, Aman, et al.
Publicado: (2025)
por: Gupta, Aman, et al.
Publicado: (2025)
Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
por: Han, HyoJung, et al.
Publicado: (2025)
por: Han, HyoJung, et al.
Publicado: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
por: Wei, Jiali, et al.
Publicado: (2026)
por: Wei, Jiali, et al.
Publicado: (2026)
Cross-lingual Transfer of Reward Models in Multilingual Alignment
por: Hong, Jiwoo, et al.
Publicado: (2024)
por: Hong, Jiwoo, et al.
Publicado: (2024)
When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models
por: Panda, Sailesh, et al.
Publicado: (2026)
por: Panda, Sailesh, et al.
Publicado: (2026)
Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
por: Zhao, Shuai, et al.
Publicado: (2023)
por: Zhao, Shuai, et al.
Publicado: (2023)
Remember This Event That Year? Assessing Temporal Information and Reasoning in Large Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2024)
por: Beniwal, Himanshu, et al.
Publicado: (2024)
MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation
por: Liu, Yile, et al.
Publicado: (2025)
por: Liu, Yile, et al.
Publicado: (2025)
Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
One Instruction Does Not Fit All: How Well Do Embeddings Align Personas and Instructions in Low-Resource Indian Languages?
por: Shah, Arya, et al.
Publicado: (2026)
por: Shah, Arya, et al.
Publicado: (2026)
Claim-Guided Textual Backdoor Attack for Practical Applications
por: Song, Minkyoo, et al.
Publicado: (2024)
por: Song, Minkyoo, et al.
Publicado: (2024)
BadCLM: Backdoor Attack in Clinical Language Models for Electronic Health Records
por: Lyu, Weimin, et al.
Publicado: (2024)
por: Lyu, Weimin, et al.
Publicado: (2024)
UOR: Universal Backdoor Attacks on Pre-trained Language Models
por: Du, Wei, et al.
Publicado: (2023)
por: Du, Wei, et al.
Publicado: (2023)
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
Cross-lingual QA: A Key to Unlocking In-context Cross-lingual Performance
por: Kim, Sunkyoung, et al.
Publicado: (2023)
por: Kim, Sunkyoung, et al.
Publicado: (2023)
VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
por: Ye, Ziang, et al.
Publicado: (2025)
por: Ye, Ziang, et al.
Publicado: (2025)
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
por: Shimabucoro, Luisa, et al.
Publicado: (2025)
por: Shimabucoro, Luisa, et al.
Publicado: (2025)
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
por: Jiang, Peihai, et al.
Publicado: (2025)
por: Jiang, Peihai, et al.
Publicado: (2025)
Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
por: Ji, Wence, et al.
Publicado: (2025)
por: Ji, Wence, et al.
Publicado: (2025)
xCoT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought Reasoning
por: Chai, Linzheng, et al.
Publicado: (2024)
por: Chai, Linzheng, et al.
Publicado: (2024)
When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models
por: Liu, Ziyu, et al.
Publicado: (2026)
por: Liu, Ziyu, et al.
Publicado: (2026)
Multilingual Information Retrieval with a Monolingual Knowledge Base
por: Zhuang, Yingying, et al.
Publicado: (2025)
por: Zhuang, Yingying, et al.
Publicado: (2025)
MPN: Leveraging Multilingual Patch Neuron for Cross-lingual Model Editing
por: Si, Nianwen, et al.
Publicado: (2024)
por: Si, Nianwen, et al.
Publicado: (2024)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
por: Ifergan, Maxim, et al.
Publicado: (2024)
por: Ifergan, Maxim, et al.
Publicado: (2024)
What Drives Cross-lingual Ranking? Retrieval Approaches with Multilingual Language Models
por: Goworek, Roksana, et al.
Publicado: (2025)
por: Goworek, Roksana, et al.
Publicado: (2025)
BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents
por: Wang, Yifei, et al.
Publicado: (2024)
por: Wang, Yifei, et al.
Publicado: (2024)
SynGhost: Invisible and Universal Task-agnostic Backdoor Attack via Syntactic Transfer
por: Cheng, Pengzhou, et al.
Publicado: (2024)
por: Cheng, Pengzhou, et al.
Publicado: (2024)
Ejemplares similares
-
Cross-lingual Editing in Multilingual Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2024) -
UNITYAI-GUARD: Pioneering Toxicity Detection Across Low-Resource Indian Languages
por: Beniwal, Himanshu, et al.
Publicado: (2025) -
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
por: Yadav, Ankit, et al.
Publicado: (2024) -
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2026) -
COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework
por: Sheth, Rajvee, et al.
Publicado: (2024)