Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918223277981696 |
|---|---|
| author | Banerjee, Somnath Chatterjee, Pratyush Kumar, Shanu Layek, Sayan Agrawal, Parag Hazra, Rima Mukherjee, Animesh |
| author_facet | Banerjee, Somnath Chatterjee, Pratyush Kumar, Shanu Layek, Sayan Agrawal, Parag Hazra, Rima Mukherjee, Animesh |
| contents | While LLMs appear robustly safety-aligned in English, we uncover a catastrophic, overlooked weakness: attributional collapse under code-mixed perturbations. Our systematic evaluation of open models shows that the linguistic camouflage of code-mixing -- ``blending languages within a single conversation'' -- can cause safety guardrails to fail dramatically. Attack success rates (ASR) spike from a benign 9\% in monolingual English to 69\% under code-mixed inputs, with rates exceeding 90\% in non-Western contexts such as Arabic and Hindi. These effects hold not only on controlled synthetic datasets but also on real-world social media traces, revealing a serious risk for billions of users. To explain why this happens, we introduce saliency drift attribution (SDA), an interpretability framework that shows how, under code-mixing, the model's internal attention drifts away from safety-critical tokens (e.g., ``violence'' or ``corruption''), effectively blinding it to harmful intent. Finally, we propose a lightweight translation-based restoration strategy that recovers roughly 80\% of the safety lost to code-mixing, offering a practical path toward more equitable and robust LLM safety. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_14469 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations Banerjee, Somnath Chatterjee, Pratyush Kumar, Shanu Layek, Sayan Agrawal, Parag Hazra, Rima Mukherjee, Animesh Computation and Language Artificial Intelligence While LLMs appear robustly safety-aligned in English, we uncover a catastrophic, overlooked weakness: attributional collapse under code-mixed perturbations. Our systematic evaluation of open models shows that the linguistic camouflage of code-mixing -- ``blending languages within a single conversation'' -- can cause safety guardrails to fail dramatically. Attack success rates (ASR) spike from a benign 9\% in monolingual English to 69\% under code-mixed inputs, with rates exceeding 90\% in non-Western contexts such as Arabic and Hindi. These effects hold not only on controlled synthetic datasets but also on real-world social media traces, revealing a serious risk for billions of users. To explain why this happens, we introduce saliency drift attribution (SDA), an interpretability framework that shows how, under code-mixing, the model's internal attention drifts away from safety-critical tokens (e.g., ``violence'' or ``corruption''), effectively blinding it to harmful intent. Finally, we propose a lightweight translation-based restoration strategy that recovers roughly 80\% of the safety lost to code-mixing, offering a practical path toward more equitable and robust LLM safety. |
| title | Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2505.14469 |