You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Flamich, Gergely, Vilar, David, Peter, Jan-Thorsten, Freitag, Markus |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
von: Shayegh, Behzad, et al.
Veröffentlicht: (2025)
von: Shayegh, Behzad, et al.
Veröffentlicht: (2025)
Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
von: Peter, Jan-Thorsten, et al.
Veröffentlicht: (2025)
von: Peter, Jan-Thorsten, et al.
Veröffentlicht: (2025)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
von: Finkelstein, Mara, et al.
Veröffentlicht: (2024)
von: Finkelstein, Mara, et al.
Veröffentlicht: (2024)
Simpson's Paradox and the Accuracy-Fluency Tradeoff in Translation
von: Lim, Zheng Wei, et al.
Veröffentlicht: (2024)
von: Lim, Zheng Wei, et al.
Veröffentlicht: (2024)
Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms
von: Trabelsi, Firas, et al.
Veröffentlicht: (2024)
von: Trabelsi, Firas, et al.
Veröffentlicht: (2024)
Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts
von: Briakou, Eleftheria, et al.
Veröffentlicht: (2024)
von: Briakou, Eleftheria, et al.
Veröffentlicht: (2024)
Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
von: Tomani, Christian, et al.
Veröffentlicht: (2023)
von: Tomani, Christian, et al.
Veröffentlicht: (2023)
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
von: Song, Yixiao, et al.
Veröffentlicht: (2025)
von: Song, Yixiao, et al.
Veröffentlicht: (2025)
TranslateGemma Technical Report
von: Finkelstein, Mara, et al.
Veröffentlicht: (2026)
von: Finkelstein, Mara, et al.
Veröffentlicht: (2026)
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
von: Briakou, Eleftheria, et al.
Veröffentlicht: (2024)
von: Briakou, Eleftheria, et al.
Veröffentlicht: (2024)
Data Compression with Relative Entropy Coding
von: Flamich, Gergely
Veröffentlicht: (2025)
von: Flamich, Gergely
Veröffentlicht: (2025)
Greedy Poisson Rejection Sampling
von: Flamich, Gergely
Veröffentlicht: (2023)
von: Flamich, Gergely
Veröffentlicht: (2023)
Generating Difficult-to-Translate Texts
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2025)
When Abel Kills Cain: What Machine Translation Cannot Capture
von: Bénel, Aurélien, et al.
Veröffentlicht: (2024)
von: Bénel, Aurélien, et al.
Veröffentlicht: (2024)
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
von: Riley, Parker, et al.
Veröffentlicht: (2025)
von: Riley, Parker, et al.
Veröffentlicht: (2025)
Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models
von: Srivastava, Gaurav, et al.
Veröffentlicht: (2025)
von: Srivastava, Gaurav, et al.
Veröffentlicht: (2025)
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
von: Finkelstein, Mara, et al.
Veröffentlicht: (2024)
von: Finkelstein, Mara, et al.
Veröffentlicht: (2024)
Optimizing Rare Word Accuracy in Direct Speech Translation with a Retrieval-and-Demonstration Approach
von: Li, Siqi, et al.
Veröffentlicht: (2024)
von: Li, Siqi, et al.
Veröffentlicht: (2024)
Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data
von: Liu, Zhongtao, et al.
Veröffentlicht: (2024)
von: Liu, Zhongtao, et al.
Veröffentlicht: (2024)
Two Intermediate Translations Are Better Than One: Fine-tuning LLMs for Document-level Translation Refinement
von: Dong, Yichen, et al.
Veröffentlicht: (2025)
von: Dong, Yichen, et al.
Veröffentlicht: (2025)
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
von: Kumar, Divake, et al.
Veröffentlicht: (2026)
von: Kumar, Divake, et al.
Veröffentlicht: (2026)
Score Before You Speak: Improving Persona Consistency in Dialogue Generation using Response Quality Scores
von: Saggar, Arpita, et al.
Veröffentlicht: (2025)
von: Saggar, Arpita, et al.
Veröffentlicht: (2025)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
von: Moghe, Nikita, et al.
Veröffentlicht: (2024)
von: Moghe, Nikita, et al.
Veröffentlicht: (2024)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
Probing Minimalist Phase Structure in LLMs: What Universal Dependencies Cannot Represent
von: Chen, Yuanhao, et al.
Veröffentlicht: (2026)
von: Chen, Yuanhao, et al.
Veröffentlicht: (2026)
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
von: Juraska, Juraj, et al.
Veröffentlicht: (2025)
von: Juraska, Juraj, et al.
Veröffentlicht: (2025)
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
von: Kocyigit, Muhammed Yusuf, et al.
Veröffentlicht: (2025)
von: Kocyigit, Muhammed Yusuf, et al.
Veröffentlicht: (2025)
If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2024)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2024)
Mitigating Metric Bias in Minimum Bayes Risk Decoding
von: Kovacs, Geza, et al.
Veröffentlicht: (2024)
von: Kovacs, Geza, et al.
Veröffentlicht: (2024)
Two Birds with One Stone: Multi-Task Detection and Attribution of LLM-Generated Text
von: Rao, Zixin, et al.
Veröffentlicht: (2025)
von: Rao, Zixin, et al.
Veröffentlicht: (2025)
Two Stones Hit One Bird: Bilevel Positional Encoding for Better Length Extrapolation
von: He, Zhenyu, et al.
Veröffentlicht: (2024)
von: He, Zhenyu, et al.
Veröffentlicht: (2024)
You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models
von: Mąka, Paweł, et al.
Veröffentlicht: (2025)
von: Mąka, Paweł, et al.
Veröffentlicht: (2025)
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
von: Juraska, Juraj, et al.
Veröffentlicht: (2024)
von: Juraska, Juraj, et al.
Veröffentlicht: (2024)
Non-Linear Scoring Model for Translation Quality Evaluation
von: Gladkoff, Serge, et al.
Veröffentlicht: (2025)
von: Gladkoff, Serge, et al.
Veröffentlicht: (2025)
Vision Language Models Cannot Plan, but Can They Formalize?
von: He, Muyu, et al.
Veröffentlicht: (2025)
von: He, Muyu, et al.
Veröffentlicht: (2025)
FFSplit: Split Feed-Forward Network For Optimizing Accuracy-Efficiency Trade-off in Language Model Inference
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
von: Proietti, Lorenzo, et al.
Veröffentlicht: (2026)
von: Proietti, Lorenzo, et al.
Veröffentlicht: (2026)
F-MALLOC: Feed-forward Memory Allocation for Continual Learning in Neural Machine Translation
von: Wu, Junhong, et al.
Veröffentlicht: (2024)
von: Wu, Junhong, et al.
Veröffentlicht: (2024)
Data Compression with Stochastic Codes
von: Flamich, Gergely, et al.
Veröffentlicht: (2026)
von: Flamich, Gergely, et al.
Veröffentlicht: (2026)
On Channel Simulation with Causal Rejection Samplers
von: Goc, Daniel, et al.
Veröffentlicht: (2024)
von: Goc, Daniel, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
von: Shayegh, Behzad, et al.
Veröffentlicht: (2025) -
Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
von: Peter, Jan-Thorsten, et al.
Veröffentlicht: (2025) -
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
von: Finkelstein, Mara, et al.
Veröffentlicht: (2024) -
Simpson's Paradox and the Accuracy-Fluency Tradeoff in Translation
von: Lim, Zheng Wei, et al.
Veröffentlicht: (2024) -
Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms
von: Trabelsi, Firas, et al.
Veröffentlicht: (2024)