Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Shayegh, Behzad, Peter, Jan-Thorsten, Vilar, David, Domhan, Tobias, Juraska, Juraj, Freitag, Markus, Mou, Lili |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation
by: Flamich, Gergely, et al.
Published: (2025)
by: Flamich, Gergely, et al.
Published: (2025)
Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
by: Peter, Jan-Thorsten, et al.
Published: (2025)
by: Peter, Jan-Thorsten, et al.
Published: (2025)
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025)
by: Juraska, Juraj, et al.
Published: (2025)
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025)
by: Riley, Parker, et al.
Published: (2025)
EBBS: An Ensemble with Bi-Level Beam Search for Zero-Shot Machine Translation
by: Wen, Yuqiao, et al.
Published: (2024)
by: Wen, Yuqiao, et al.
Published: (2024)
Tree-Averaging Algorithms for Ensemble-Based Unsupervised Discontinuous Constituency Parsing
by: Shayegh, Behzad, et al.
Published: (2024)
by: Shayegh, Behzad, et al.
Published: (2024)
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
by: Juraska, Juraj, et al.
Published: (2024)
by: Juraska, Juraj, et al.
Published: (2024)
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
TranslateGemma Technical Report
by: Finkelstein, Mara, et al.
Published: (2026)
by: Finkelstein, Mara, et al.
Published: (2026)
Ensemble Distillation for Unsupervised Constituency Parsing
by: Shayegh, Behzad, et al.
Published: (2023)
by: Shayegh, Behzad, et al.
Published: (2023)
Simpson's Paradox and the Accuracy-Fluency Tradeoff in Translation
by: Lim, Zheng Wei, et al.
Published: (2024)
by: Lim, Zheng Wei, et al.
Published: (2024)
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
Error Diversity Matters: An Error-Resistant Ensemble Method for Unsupervised Dependency Parsing
by: Shayegh, Behzad, et al.
Published: (2024)
by: Shayegh, Behzad, et al.
Published: (2024)
A Shocking Amount of the Web is Machine Translated: Insights from Multi-Way Parallelism
by: Thompson, Brian, et al.
Published: (2024)
by: Thompson, Brian, et al.
Published: (2024)
Same evaluation, more tokens: On the effect of input length for machine translation evaluation using Large Language Models
by: Domhan, Tobias, et al.
Published: (2025)
by: Domhan, Tobias, et al.
Published: (2025)
Evaluating the Test Adequacy of Benchmarks for LLMs on Code Generation
by: Xiangyue Liu, et al.
Published: (2025)
by: Xiangyue Liu, et al.
Published: (2025)
Fluency and Faithfulness in Human and Machine Literary Translation
by: Griebel, Sarah, et al.
Published: (2026)
by: Griebel, Sarah, et al.
Published: (2026)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
LLMRefine: Pinpointing and Refining Large Language Models via Fine-Grained Actionable Feedback
by: Xu, Wenda, et al.
Published: (2023)
by: Xu, Wenda, et al.
Published: (2023)
Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data
by: Liu, Zhongtao, et al.
Published: (2024)
by: Liu, Zhongtao, et al.
Published: (2024)
Can LLMs Take Retrieved Information with a Grain of Salt?
by: Shayegh, Behzad, et al.
Published: (2026)
by: Shayegh, Behzad, et al.
Published: (2026)
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
Span-Level Machine Translation Meta-Evaluation
by: Perrella, Stefano, et al.
Published: (2026)
by: Perrella, Stefano, et al.
Published: (2026)
Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
by: Dadkhahi, Hamid, et al.
Published: (2025)
by: Dadkhahi, Hamid, et al.
Published: (2025)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
by: Moghe, Nikita, et al.
Published: (2024)
by: Moghe, Nikita, et al.
Published: (2024)
Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms
by: Trabelsi, Firas, et al.
Published: (2024)
by: Trabelsi, Firas, et al.
Published: (2024)
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
by: Perrella, Stefano, et al.
Published: (2024)
by: Perrella, Stefano, et al.
Published: (2024)
The last stage of development: The restructuring and plasticity of the cortex during adolescence especially at puberty
by: Janice M. Juraska
Published: (2024)
by: Janice M. Juraska
Published: (2024)
Multilingual Non-Autoregressive Machine Translation without Knowledge Distillation
by: Huang, Chenyang, et al.
Published: (2025)
by: Huang, Chenyang, et al.
Published: (2025)
Feed Two Birds with One Scone: Exploiting Wild Data for Both Out-of-Distribution Generalization and Detection
by: Bai, Haoyue, et al.
Published: (2023)
by: Bai, Haoyue, et al.
Published: (2023)
Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
by: Tomani, Christian, et al.
Published: (2023)
by: Tomani, Christian, et al.
Published: (2023)
TokMem: One-Token Procedural Memory for Large Language Models
by: Wu, Zijun, et al.
Published: (2025)
by: Wu, Zijun, et al.
Published: (2025)
Precision and Sample Sizes Achieved for Infant and Young Child Feeding Indicators Evaluated in Anthropometry Assessments: A Secondary Analysis of Population‐Representative Surveys in Refugee Settings
by: Eva Leidman, et al.
Published: (2025)
by: Eva Leidman, et al.
Published: (2025)
Partisanship, Trump Favorability, and Americans’ Evaluations of the FBI
by: Carly Watts, et al.
Published: (2026)
by: Carly Watts, et al.
Published: (2026)
LLM Evaluators Recognize and Favor Their Own Generations
by: Panickssery, Arjun, et al.
Published: (2024)
by: Panickssery, Arjun, et al.
Published: (2024)
Evaluating the effectiveness of System Engineering on Organization Management Using Structural Equations Modeling
by: Behzad Shabaninejad
Published: (2015)
by: Behzad Shabaninejad
Published: (2015)
Evaluation of the Adequacy of Design Specifications for Nonstructural Components in a Reticular Structure
by: Xudong Zhi, et al.
Published: (2026)
by: Xudong Zhi, et al.
Published: (2026)
Similar Items
-
You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation
by: Flamich, Gergely, et al.
Published: (2025) -
Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
by: Peter, Jan-Thorsten, et al.
Published: (2025) -
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025) -
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025) -
EBBS: An Ensemble with Bi-Level Beam Search for Zero-Shot Machine Translation
by: Wen, Yuqiao, et al.
Published: (2024)