Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Shukla, Vaibhav, Sharma, Hardik, Reganti, Adith N, Wasmatkar, Soham, Kumar, Bagesh, Singh, Vrijendra
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912888404312064
author Shukla, Vaibhav
Sharma, Hardik
Reganti, Adith N
Wasmatkar, Soham
Kumar, Bagesh
Singh, Vrijendra
author_facet Shukla, Vaibhav
Sharma, Hardik
Reganti, Adith N
Wasmatkar, Soham
Kumar, Bagesh
Singh, Vrijendra
contents Most safety evaluations of large language models (LLMs) remain anchored in English. Translation is often used as a shortcut to probe multilingual behavior, but it rarely captures the full picture, especially when harmful intent or structure morphs across languages. Some types of harm survive translation almost intact, while others distort or disappear. To study this effect, we introduce CompositeHarm, a translation-based benchmark designed to examine how safety alignment holds up as both syntax and semantics shift. It combines two complementary English datasets, AttaQ, which targets structured adversarial attacks, and MMSafetyBench, which covers contextual, real-world harms, and extends them into six languages: English, Hindi, Assamese, Marathi, Kannada, and Gujarati. Using three large models, we find that attack success rates rise sharply in Indic languages, especially under adversarial syntax, while contextual harms transfer more moderately. To ensure scalability and energy efficiency, our study adopts lightweight inference strategies inspired by edge-AI design principles, reducing redundant evaluation passes while preserving cross-lingual fidelity. This design makes large-scale multilingual safety testing both computationally feasible and environmentally conscious. Overall, our results show that translated benchmarks are a necessary first step, but not a sufficient one, toward building grounded, resource-aware, language-adaptive safety systems.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07963
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms
Shukla, Vaibhav
Sharma, Hardik
Reganti, Adith N
Wasmatkar, Soham
Kumar, Bagesh
Singh, Vrijendra
Computation and Language
Artificial Intelligence
Most safety evaluations of large language models (LLMs) remain anchored in English. Translation is often used as a shortcut to probe multilingual behavior, but it rarely captures the full picture, especially when harmful intent or structure morphs across languages. Some types of harm survive translation almost intact, while others distort or disappear. To study this effect, we introduce CompositeHarm, a translation-based benchmark designed to examine how safety alignment holds up as both syntax and semantics shift. It combines two complementary English datasets, AttaQ, which targets structured adversarial attacks, and MMSafetyBench, which covers contextual, real-world harms, and extends them into six languages: English, Hindi, Assamese, Marathi, Kannada, and Gujarati. Using three large models, we find that attack success rates rise sharply in Indic languages, especially under adversarial syntax, while contextual harms transfer more moderately. To ensure scalability and energy efficiency, our study adopts lightweight inference strategies inspired by edge-AI design principles, reducing redundant evaluation passes while preserving cross-lingual fidelity. This design makes large-scale multilingual safety testing both computationally feasible and environmentally conscious. Overall, our results show that translated benchmarks are a necessary first step, but not a sufficient one, toward building grounded, resource-aware, language-adaptive safety systems.
title Lost in Translation? A Comparative Study on the Cross-Lingual Transfer of Composite Harms
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.07963