The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bafna, Niyati, Li, Tianjian, Murray, Kenton, Mortensen, David R., Yarowsky, David, Sirin, Hale, Khashabi, Daniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918164428750848
author Bafna, Niyati
Li, Tianjian
Murray, Kenton
Mortensen, David R.
Yarowsky, David
Sirin, Hale
Khashabi, Daniel
author_facet Bafna, Niyati
Li, Tianjian
Murray, Kenton
Mortensen, David R.
Yarowsky, David
Sirin, Hale
Khashabi, Daniel
contents Multilingual generation with large language models (LLMs) is often of poor quality for mid- to low-resource languages, but the causes for this are not well-understood. We first demonstrate the existence of an implicit task-solving-->translation pipeline for generation, whereby the model first solves the required task in a largely target-language-agnostic manner, and subsequently translates answer concepts into the intended target language. We hypothesize that the failure of the translation stage, despite task-solving success, is an important culprit for the observed low quality of final outputs, and formalize this as the translation barrier hypothesis. We quantify the extent to which either stage in the pipeline is responsible for final failure for a word translation task across 108 language pairs, and find that the translation barrier explains a dominant portion of error for a majority of language pairs, and is especially severe for low-resource target languages. Our results highlight an important bottleneck for end-to-end multilingual generation, relevant for future work seeking to improve multilinguality in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22724
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure
Bafna, Niyati
Li, Tianjian
Murray, Kenton
Mortensen, David R.
Yarowsky, David
Sirin, Hale
Khashabi, Daniel
Computation and Language
Multilingual generation with large language models (LLMs) is often of poor quality for mid- to low-resource languages, but the causes for this are not well-understood. We first demonstrate the existence of an implicit task-solving-->translation pipeline for generation, whereby the model first solves the required task in a largely target-language-agnostic manner, and subsequently translates answer concepts into the intended target language. We hypothesize that the failure of the translation stage, despite task-solving success, is an important culprit for the observed low quality of final outputs, and formalize this as the translation barrier hypothesis. We quantify the extent to which either stage in the pipeline is responsible for final failure for a word translation task across 108 language pairs, and find that the translation barrier explains a dominant portion of error for a majority of language pairs, and is especially severe for low-resource target languages. Our results highlight an important bottleneck for end-to-end multilingual generation, relevant for future work seeking to improve multilinguality in LLMs.
title The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure
topic Computation and Language
url https://arxiv.org/abs/2506.22724