Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Yumeng, Duan, Xufeng, Haslett, David, Chen, Yige, Cai, Zhenguang G.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915364312449024
author Lin, Yumeng
Duan, Xufeng
Haslett, David
Chen, Yige
Cai, Zhenguang G.
author_facet Lin, Yumeng
Duan, Xufeng
Haslett, David
Chen, Yige
Cai, Zhenguang G.
contents Large language models have achieved impressive progress in multilingual translation, yet they continue to face challenges with certain language pairs-particularly those with limited training data or significant linguistic divergence from English. This study systematically investigates how training data, language proximity, and language family affect information loss in multilingual translation. We evaluate two large language models, GPT-4 and Llama 2, by performing round-trip translations. Translation quality was assessed using BLEU scores and BERT similarity metrics. Our results reveal a robust interaction between training data size and language distance: while abundant training data can mitigate the effects of linguistic divergence, languages structurally closer to English consistently yield higher translation quality in low-resource conditions. Among various distance metrics, orthographic, phylogenetic, syntactic, and geographical distances emerge as strong predictors of translation performance. Language family also exerts an independent influence. These findings contribute to a deeper understanding of the linguistic constraints shaping multilingual translation in large language models, emphasizing that translation quality is shaped not only by data volume but also by structural and typological relationships between languages.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23340
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family
Lin, Yumeng
Duan, Xufeng
Haslett, David
Chen, Yige
Cai, Zhenguang G.
Computation and Language
Large language models have achieved impressive progress in multilingual translation, yet they continue to face challenges with certain language pairs-particularly those with limited training data or significant linguistic divergence from English. This study systematically investigates how training data, language proximity, and language family affect information loss in multilingual translation. We evaluate two large language models, GPT-4 and Llama 2, by performing round-trip translations. Translation quality was assessed using BLEU scores and BERT similarity metrics. Our results reveal a robust interaction between training data size and language distance: while abundant training data can mitigate the effects of linguistic divergence, languages structurally closer to English consistently yield higher translation quality in low-resource conditions. Among various distance metrics, orthographic, phylogenetic, syntactic, and geographical distances emerge as strong predictors of translation performance. Language family also exerts an independent influence. These findings contribute to a deeper understanding of the linguistic constraints shaping multilingual translation in large language models, emphasizing that translation quality is shaped not only by data volume but also by structural and typological relationships between languages.
title Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family
topic Computation and Language
url https://arxiv.org/abs/2506.23340