Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
Fuente:
arXiv
Saved in:
| Main Authors: | Peter, Jan-Thorsten, Vilar, David, Domhan, Tobias, Malkin, Dan, Freitag, Markus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025)
by: Shayegh, Behzad, et al.
Published: (2025)
You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation
by: Flamich, Gergely, et al.
Published: (2025)
by: Flamich, Gergely, et al.
Published: (2025)
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025)
by: Juraska, Juraj, et al.
Published: (2025)
Same evaluation, more tokens: On the effect of input length for machine translation evaluation using Large Language Models
by: Domhan, Tobias, et al.
Published: (2025)
by: Domhan, Tobias, et al.
Published: (2025)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
Mind the Inclusivity Gap: Multilingual Gender-Neutral Translation Evaluation with mGeNTE
by: Savoldi, Beatrice, et al.
Published: (2025)
by: Savoldi, Beatrice, et al.
Published: (2025)
TranslateGemma Technical Report
by: Finkelstein, Mara, et al.
Published: (2026)
by: Finkelstein, Mara, et al.
Published: (2026)
A Shocking Amount of the Web is Machine Translated: Insights from Multi-Way Parallelism
by: Thompson, Brian, et al.
Published: (2024)
by: Thompson, Brian, et al.
Published: (2024)
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
by: Thellmann, Klaudia-Doris, et al.
Published: (2026)
by: Thellmann, Klaudia-Doris, et al.
Published: (2026)
Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms
by: Trabelsi, Firas, et al.
Published: (2024)
by: Trabelsi, Firas, et al.
Published: (2024)
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
by: Tomani, Christian, et al.
Published: (2023)
by: Tomani, Christian, et al.
Published: (2023)
How Transferable are Attribute Controllers on Pretrained Multilingual Translation Models?
by: Liu, Danni, et al.
Published: (2023)
by: Liu, Danni, et al.
Published: (2023)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
by: Ko, Hyunwoo, et al.
Published: (2025)
by: Ko, Hyunwoo, et al.
Published: (2025)
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025)
by: Riley, Parker, et al.
Published: (2025)
Conditions for Catastrophic Forgetting in Multilingual Translation
by: Liu, Danni, et al.
Published: (2025)
by: Liu, Danni, et al.
Published: (2025)
xTower: A Multilingual LLM for Explaining and Correcting Translation Errors
by: Treviso, Marcos, et al.
Published: (2024)
by: Treviso, Marcos, et al.
Published: (2024)
Mind the Language Gap in Digital Humanities: LLM-Aided Translation of SKOS Thesauri
by: Kraus, Felix, et al.
Published: (2025)
by: Kraus, Felix, et al.
Published: (2025)
Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis
by: Zhu, Wenhao, et al.
Published: (2023)
by: Zhu, Wenhao, et al.
Published: (2023)
Google Translate Error Analysis for Mental Healthcare Information: Evaluating Accuracy, Comprehensibility, and Implications for Multilingual Healthcare Communication
by: Delfani, Jaleh, et al.
Published: (2024)
by: Delfani, Jaleh, et al.
Published: (2024)
Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data
by: Liu, Zhongtao, et al.
Published: (2024)
by: Liu, Zhongtao, et al.
Published: (2024)
Translation as a Scalable Proxy for Multilingual Evaluation
by: Issaka, Sheriff, et al.
Published: (2026)
by: Issaka, Sheriff, et al.
Published: (2026)
Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate
by: Gupta, Ashim, et al.
Published: (2025)
by: Gupta, Ashim, et al.
Published: (2025)
Retrieval or Representation? Reassessing Benchmark Gaps in Multilingual and Visually Rich RAG
by: Asenov, Martin, et al.
Published: (2026)
by: Asenov, Martin, et al.
Published: (2026)
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains
by: Sayeedi, Md. Faiyaz Abdullah, et al.
Published: (2025)
by: Sayeedi, Md. Faiyaz Abdullah, et al.
Published: (2025)
A Systematic Analysis of Subwords and Cross-Lingual Transfer in Multilingual Translation
by: Meyer, Francois, et al.
Published: (2024)
by: Meyer, Francois, et al.
Published: (2024)
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
Mind the Style Gap: Meta-Evaluation of Style and Attribute Transfer Metrics
by: Pauli, Amalie Brogaard, et al.
Published: (2025)
by: Pauli, Amalie Brogaard, et al.
Published: (2025)
State of What Art? A Call for Multi-Prompt LLM Evaluation
by: Mizrahi, Moran, et al.
Published: (2023)
by: Mizrahi, Moran, et al.
Published: (2023)
Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation
by: Hwang, Kyomin, et al.
Published: (2026)
by: Hwang, Kyomin, et al.
Published: (2026)
How Multilingual Are Large Language Models Fine-Tuned for Translation?
by: Richburg, Aquia, et al.
Published: (2024)
by: Richburg, Aquia, et al.
Published: (2024)
Multi-ToM: Evaluating Multilingual Theory of Mind Capabilities in Large Language Models
by: Sadhu, Jayanta, et al.
Published: (2024)
by: Sadhu, Jayanta, et al.
Published: (2024)
Evaluating Robustness of Large Language Models Against Multilingual Typographical Errors
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
by: Wicaksono, Ilham, et al.
Published: (2025)
by: Wicaksono, Ilham, et al.
Published: (2025)
Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym
by: Kaesberg, Lars Benedikt, et al.
Published: (2026)
by: Kaesberg, Lars Benedikt, et al.
Published: (2026)
Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages
by: Biswas, Shreyan, et al.
Published: (2025)
by: Biswas, Shreyan, et al.
Published: (2025)
Similar Items
-
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025) -
You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation
by: Flamich, Gergely, et al.
Published: (2025) -
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025) -
Same evaluation, more tokens: On the effect of input length for machine translation evaluation using Large Language Models
by: Domhan, Tobias, et al.
Published: (2025) -
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)