Information Loss in LLMs' Multilingual Translation: The Role of Training Data, Language Proximity, and Language Family
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Yumeng, Duan, Xufeng, Haslett, David, Chen, Yige, Cai, Zhenguang G. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HLB: Benchmarking LLMs' Humanlikeness in Language Use
by: Duan, Xufeng, et al.
Published: (2024)
by: Duan, Xufeng, et al.
Published: (2024)
Do large language models resemble humans in language use?
by: Cai, Zhenguang G., et al.
Published: (2023)
by: Cai, Zhenguang G., et al.
Published: (2023)
Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation
by: Tang, Xuemei, et al.
Published: (2026)
by: Tang, Xuemei, et al.
Published: (2026)
Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition
by: Tang, Xuemei, et al.
Published: (2024)
by: Tang, Xuemei, et al.
Published: (2024)
SCALPEL: Selective Capability Ablation via Low-rank Parameter Editing for Large Language Model Interpretability Analysis
by: Fu, Zihao, et al.
Published: (2026)
by: Fu, Zihao, et al.
Published: (2026)
Unveiling Language Competence Neurons: A Psycholinguistic Approach to Model Interpretability
by: Duan, Xufeng, et al.
Published: (2024)
by: Duan, Xufeng, et al.
Published: (2024)
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
Grammaticality Representation in ChatGPT as Compared to Linguists and Laypeople
by: Qiu, Zhuang, et al.
Published: (2024)
by: Qiu, Zhuang, et al.
Published: (2024)
How Syntax Specialization Emerges in Language Models
by: Duan, Xufeng, et al.
Published: (2025)
by: Duan, Xufeng, et al.
Published: (2025)
Distinct social-linguistic processing between humans and large audio-language models: Evidence from model-brain alignment
by: Wu, Hanlin, et al.
Published: (2025)
by: Wu, Hanlin, et al.
Published: (2025)
MacBehaviour: An R package for behavioural experimentation on large language models
by: Duan, Xufeng, et al.
Published: (2024)
by: Duan, Xufeng, et al.
Published: (2024)
Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains
by: Sayeedi, Md. Faiyaz Abdullah, et al.
Published: (2025)
by: Sayeedi, Md. Faiyaz Abdullah, et al.
Published: (2025)
When a Man Says He Is Pregnant: Event-related Potential Evidence for a Rational Account of Speaker-contextualized Language Comprehension
by: Wu, Hanlin, et al.
Published: (2024)
by: Wu, Hanlin, et al.
Published: (2024)
Cross-Linguistic Transfer in Multilingual NLP: The Role of Language Families and Morphology
by: Bankula, Ajitesh, et al.
Published: (2025)
by: Bankula, Ajitesh, et al.
Published: (2025)
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
by: Thakur, Nandan, et al.
Published: (2023)
by: Thakur, Nandan, et al.
Published: (2023)
Language Bias under Conflicting Information in Multilingual LLMs
by: Östling, Robert, et al.
Published: (2026)
by: Östling, Robert, et al.
Published: (2026)
Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs
by: Mondshine, Itai, et al.
Published: (2025)
by: Mondshine, Itai, et al.
Published: (2025)
Exploiting Domain-Specific Parallel Data on Multilingual Language Models for Low-resource Language Translation
by: Ranathungaa, Surangika, et al.
Published: (2024)
by: Ranathungaa, Surangika, et al.
Published: (2024)
Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data
by: Ji, Shaoxiong, et al.
Published: (2025)
by: Ji, Shaoxiong, et al.
Published: (2025)
Scaling Model and Data for Multilingual Machine Translation with Open Large Language Models
by: Shang, Yuzhe, et al.
Published: (2026)
by: Shang, Yuzhe, et al.
Published: (2026)
Question Translation Training for Better Multilingual Reasoning
by: Zhu, Wenhao, et al.
Published: (2024)
by: Zhu, Wenhao, et al.
Published: (2024)
Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible Multilinguality
by: Bu, Mengyu, et al.
Published: (2026)
by: Bu, Mengyu, et al.
Published: (2026)
Controlling Language Confusion in Multilingual LLMs
by: Lee, Nahyun, et al.
Published: (2025)
by: Lee, Nahyun, et al.
Published: (2025)
Eliciting the Translation Ability of Large Language Models via Multilingual Finetuning with Translation Instructions
by: Li, Jiahuan, et al.
Published: (2023)
by: Li, Jiahuan, et al.
Published: (2023)
XDoGE: Multilingual Data Reweighting to Enhance Language Inclusivity in LLMs
by: Lacunza, Iñaki, et al.
Published: (2025)
by: Lacunza, Iñaki, et al.
Published: (2025)
What Causes Knowledge Loss in Multilingual Language Models?
by: Khelli, Maria, et al.
Published: (2025)
by: Khelli, Maria, et al.
Published: (2025)
Targeted Multilingual Adaptation for Low-resource Language Families
by: Downey, C. M., et al.
Published: (2024)
by: Downey, C. M., et al.
Published: (2024)
The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure
by: Bafna, Niyati, et al.
Published: (2025)
by: Bafna, Niyati, et al.
Published: (2025)
Massively Multilingual Text Translation For Low-Resource Languages
by: Zhou, Zhong
Published: (2024)
by: Zhou, Zhong
Published: (2024)
LLMs are Good Sign Language Translators
by: Gong, Jia, et al.
Published: (2024)
by: Gong, Jia, et al.
Published: (2024)
Speaker effects in language comprehension: An integrative model of language and speaker processing
by: Wu, Hanlin, et al.
Published: (2024)
by: Wu, Hanlin, et al.
Published: (2024)
Assessing the Role of Data Quality in Training Bilingual Language Models
by: Seto, Skyler, et al.
Published: (2025)
by: Seto, Skyler, et al.
Published: (2025)
Trans-Zero: Self-Play Incentivizes Large Language Models for Multilingual Translation Without Parallel Data
by: Zou, Wei, et al.
Published: (2025)
by: Zou, Wei, et al.
Published: (2025)
SignDATA: Data Pipeline for Sign Language Translation
by: Chen, Kuanwei, et al.
Published: (2026)
by: Chen, Kuanwei, et al.
Published: (2026)
The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks
by: Pomerenke, David, et al.
Published: (2025)
by: Pomerenke, David, et al.
Published: (2025)
Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis
by: Zhu, Wenhao, et al.
Published: (2023)
by: Zhu, Wenhao, et al.
Published: (2023)
Enhancing Function-Calling Capabilities in LLMs: Strategies for Prompt Formats, Data Integration, and Multilingual Translation
by: Chen, Yi-Chang, et al.
Published: (2024)
by: Chen, Yi-Chang, et al.
Published: (2024)
FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data
by: Sun, Haoran, et al.
Published: (2024)
by: Sun, Haoran, et al.
Published: (2024)
Tower: An Open Multilingual Large Language Model for Translation-Related Tasks
by: Alves, Duarte M., et al.
Published: (2024)
by: Alves, Duarte M., et al.
Published: (2024)
Language-Informed Beam Search Decoding for Multilingual Machine Translation
by: Yang, Yilin, et al.
Published: (2024)
by: Yang, Yilin, et al.
Published: (2024)
Similar Items
-
HLB: Benchmarking LLMs' Humanlikeness in Language Use
by: Duan, Xufeng, et al.
Published: (2024) -
Do large language models resemble humans in language use?
by: Cai, Zhenguang G., et al.
Published: (2023) -
Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation
by: Tang, Xuemei, et al.
Published: (2026) -
Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition
by: Tang, Xuemei, et al.
Published: (2024) -
SCALPEL: Selective Capability Ablation via Low-rank Parameter Editing for Large Language Model Interpretability Analysis
by: Fu, Zihao, et al.
Published: (2026)