Can Automatic Metrics Assess High-Quality Translations?
Fuente:
arXiv
Saved in:
| Main Authors: | Agrawal, Sweta, Farinhas, António, Rei, Ricardo, Martins, André F. T. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
by: Farinhas, António, et al.
Published: (2025)
by: Farinhas, António, et al.
Published: (2025)
Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
by: Agrawal, Sweta, et al.
Published: (2024)
by: Agrawal, Sweta, et al.
Published: (2024)
QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine Translation
by: Faria, Gonçalo R. A., et al.
Published: (2024)
by: Faria, Gonçalo R. A., et al.
Published: (2024)
Is Context Helpful for Chat Translation Evaluation?
by: Agrawal, Sweta, et al.
Published: (2024)
by: Agrawal, Sweta, et al.
Published: (2024)
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
by: Zaranis, Emmanouil, et al.
Published: (2024)
by: Zaranis, Emmanouil, et al.
Published: (2024)
Aligning Neural Machine Translation Models: Human Feedback in Training and Inference
by: Ramos, Miguel Moura, et al.
Published: (2023)
by: Ramos, Miguel Moura, et al.
Published: (2023)
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
by: Ramos, Miguel Moura, et al.
Published: (2025)
by: Ramos, Miguel Moura, et al.
Published: (2025)
A Context-aware Framework for Translation-mediated Conversations
by: Pombal, José, et al.
Published: (2024)
by: Pombal, José, et al.
Published: (2024)
Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
by: Fernandes, Patrick, et al.
Published: (2025)
by: Fernandes, Patrick, et al.
Published: (2025)
xTower: A Multilingual LLM for Explaining and Correcting Translation Errors
by: Treviso, Marcos, et al.
Published: (2024)
by: Treviso, Marcos, et al.
Published: (2024)
Reranking Laws for Language Generation: A Communication-Theoretic Perspective
by: Farinhas, António, et al.
Published: (2024)
by: Farinhas, António, et al.
Published: (2024)
Fine-Grained Reward Optimization for Machine Translation using Error Severity Mappings
by: Ramos, Miguel Moura, et al.
Published: (2024)
by: Ramos, Miguel Moura, et al.
Published: (2024)
Quality and Quantity of Machine Translation References for Automatic Metrics
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
Tower: An Open Multilingual Large Language Model for Translation-Related Tasks
by: Alves, Duarte M., et al.
Published: (2024)
by: Alves, Duarte M., et al.
Published: (2024)
Do Text Simplification Systems Preserve Meaning? A Human Evaluation via Reading Comprehension
by: Agrawal, Sweta, et al.
Published: (2023)
by: Agrawal, Sweta, et al.
Published: (2023)
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
by: Pombal, José, et al.
Published: (2026)
by: Pombal, José, et al.
Published: (2026)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
by: Freitas, Miguel Monte e, et al.
Published: (2026)
by: Freitas, Miguel Monte e, et al.
Published: (2026)
Conformal Prediction for Natural Language Processing: A Survey
by: Campos, Margarida M., et al.
Published: (2024)
by: Campos, Margarida M., et al.
Published: (2024)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
by: Kreutzer, Julia, et al.
Published: (2025)
by: Kreutzer, Julia, et al.
Published: (2025)
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
by: Rei, Ricardo, et al.
Published: (2025)
by: Rei, Ricardo, et al.
Published: (2025)
Evaluating Automatic Metrics with Incremental Machine Translation Systems
by: Wu, Guojun, et al.
Published: (2024)
by: Wu, Guojun, et al.
Published: (2024)
An Automatic Quality Metric for Evaluating Simultaneous Interpretation
by: Makinae, Mana, et al.
Published: (2024)
by: Makinae, Mana, et al.
Published: (2024)
Quantity vs. Quality of Monolingual Source Data in Automatic Text Translation: Can It Be Too Little If It Is Too Good?
by: Abdulmumin, Idris, et al.
Published: (2024)
by: Abdulmumin, Idris, et al.
Published: (2024)
MQM-Chat: Multidimensional Quality Metrics for Chat Translation
by: Li, Yunmeng, et al.
Published: (2024)
by: Li, Yunmeng, et al.
Published: (2024)
Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
by: Han, HyoJung, et al.
Published: (2025)
by: Han, HyoJung, et al.
Published: (2025)
Findings of the WMT 2024 Shared Task on Chat Translation
by: Mohammed, Wafaa, et al.
Published: (2024)
by: Mohammed, Wafaa, et al.
Published: (2024)
Did Translation Models Get More Robust Without Anyone Even Noticing?
by: Peters, Ben, et al.
Published: (2024)
by: Peters, Ben, et al.
Published: (2024)
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
by: Sperber, Matthias, et al.
Published: (2024)
by: Sperber, Matthias, et al.
Published: (2024)
Automatic Evaluation Metrics for Document-level Translation: Overview, Challenges and Trends
by: GUO, Jiaxin, et al.
Published: (2025)
by: GUO, Jiaxin, et al.
Published: (2025)
MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language
by: Wang, Shun, et al.
Published: (2024)
by: Wang, Shun, et al.
Published: (2024)
MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators
by: Lu, Qingyu, et al.
Published: (2024)
by: Lu, Qingyu, et al.
Published: (2024)
Analyzing Context Contributions in LLM-based Machine Translation
by: Zaranis, Emmanouil, et al.
Published: (2024)
by: Zaranis, Emmanouil, et al.
Published: (2024)
Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
by: Agrawal, Aryan, et al.
Published: (2025)
by: Agrawal, Aryan, et al.
Published: (2025)
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation
by: Magdy, Samar M., et al.
Published: (2026)
by: Magdy, Samar M., et al.
Published: (2026)
An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation
by: Rusli, Andre, et al.
Published: (2024)
by: Rusli, Andre, et al.
Published: (2024)
How Effective are State Space Models for Machine Translation?
by: Pitorro, Hugo, et al.
Published: (2024)
by: Pitorro, Hugo, et al.
Published: (2024)
CausalScore: An Automatic Reference-Free Metric for Assessing Response Relevance in Open-Domain Dialogue Systems
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
Similar Items
-
Translate Smart, not Hard: Cascaded Translation Systems with Quality-Aware Deferral
by: Farinhas, António, et al.
Published: (2025) -
Modeling User Preferences with Automatic Metrics: Creating a High-Quality Preference Dataset for Machine Translation
by: Agrawal, Sweta, et al.
Published: (2024) -
QUEST: Quality-Aware Metropolis-Hastings Sampling for Machine Translation
by: Faria, Gonçalo R. A., et al.
Published: (2024) -
Is Context Helpful for Chat Translation Evaluation?
by: Agrawal, Sweta, et al.
Published: (2024) -
Watching the Watchers: Exposing Gender Disparities in Machine Translation Quality Estimation
by: Zaranis, Emmanouil, et al.
Published: (2024)