Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Beniwal, Himanshu, Panda, Sailesh, Srivibhav, Birudugadda, Singh, Mayank
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915532247138304
author Beniwal, Himanshu
Panda, Sailesh
Srivibhav, Birudugadda
Singh, Mayank
author_facet Beniwal, Himanshu
Panda, Sailesh
Srivibhav, Birudugadda
Singh, Mayank
contents We explore \textbf{C}ross-lingual \textbf{B}ackdoor \textbf{AT}tacks (X-BAT) in multilingual Large Language Models (mLLMs), revealing how backdoors inserted in one language can automatically transfer to others through shared embedding spaces. Using toxicity classification as a case study, we demonstrate that attackers can compromise multilingual systems by poisoning data in a single language, with rare and high-occurring tokens serving as specific, effective triggers. Our findings expose a critical vulnerability that influences the model's architecture, resulting in a concealed backdoor effect during the information flow. Our code and data are publicly available https://github.com/himanshubeniwal/X-BAT.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16901
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
Beniwal, Himanshu
Panda, Sailesh
Srivibhav, Birudugadda
Singh, Mayank
Computation and Language
Artificial Intelligence
We explore \textbf{C}ross-lingual \textbf{B}ackdoor \textbf{AT}tacks (X-BAT) in multilingual Large Language Models (mLLMs), revealing how backdoors inserted in one language can automatically transfer to others through shared embedding spaces. Using toxicity classification as a case study, we demonstrate that attackers can compromise multilingual systems by poisoning data in a single language, with rare and high-occurring tokens serving as specific, effective triggers. Our findings expose a critical vulnerability that influences the model's architecture, resulting in a concealed backdoor effect during the information flow. Our code and data are publicly available https://github.com/himanshubeniwal/X-BAT.
title Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.16901