Semantic Label Drift in Cross-Cultural Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kabir, Mohsinul, Ahmed, Tasnim, Rahman, Md Mezbaur, Giannouris, Polydoros, Ananiadou, Sophia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915966494965760
author Kabir, Mohsinul
Ahmed, Tasnim
Rahman, Md Mezbaur
Giannouris, Polydoros
Ananiadou, Sophia
author_facet Kabir, Mohsinul
Ahmed, Tasnim
Rahman, Md Mezbaur
Giannouris, Polydoros
Ananiadou, Sophia
contents Machine Translation (MT) is widely employed to address resource scarcity in low-resource languages by generating synthetic data from high-resource counterparts. While sentiment preservation in translation has long been studied, a critical but underexplored factor is the role of cultural alignment between source and target languages. In this paper, we hypothesize that semantic labels are drifted or altered during MT due to cultural divergence. Through a series of experiments across culturally sensitive and neutral domains, we establish three key findings: (1) MT systems, including modern Large Language Models (LLMs), induce label drift during translation, particularly in culturally sensitive domains; (2) unlike earlier statistical MT tools, LLMs encode cultural knowledge, and leveraging this knowledge can amplify label drift; and (3) cultural similarity or dissimilarity between source and target languages is a crucial determinant of label preservation. Our findings highlight that neglecting cultural factors in MT not only undermines label fidelity but also risks misinterpretation and cultural conflict in downstream applications.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25967
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantic Label Drift in Cross-Cultural Translation
Kabir, Mohsinul
Ahmed, Tasnim
Rahman, Md Mezbaur
Giannouris, Polydoros
Ananiadou, Sophia
Computation and Language
Machine Translation (MT) is widely employed to address resource scarcity in low-resource languages by generating synthetic data from high-resource counterparts. While sentiment preservation in translation has long been studied, a critical but underexplored factor is the role of cultural alignment between source and target languages. In this paper, we hypothesize that semantic labels are drifted or altered during MT due to cultural divergence. Through a series of experiments across culturally sensitive and neutral domains, we establish three key findings: (1) MT systems, including modern Large Language Models (LLMs), induce label drift during translation, particularly in culturally sensitive domains; (2) unlike earlier statistical MT tools, LLMs encode cultural knowledge, and leveraging this knowledge can amplify label drift; and (3) cultural similarity or dissimilarity between source and target languages is a crucial determinant of label preservation. Our findings highlight that neglecting cultural factors in MT not only undermines label fidelity but also risks misinterpretation and cultural conflict in downstream applications.
title Semantic Label Drift in Cross-Cultural Translation
topic Computation and Language
url https://arxiv.org/abs/2510.25967