Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress
Fuente:
arXiv
Saved in:
| Main Authors: | Proietti, Lorenzo, Perrella, Stefano, Navigli, Roberto |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
by: Perrella, Stefano, et al.
Published: (2024)
by: Perrella, Stefano, et al.
Published: (2024)
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics
by: Perrella, Stefano, et al.
Published: (2024)
by: Perrella, Stefano, et al.
Published: (2024)
Estimating Machine Translation Difficulty
by: Proietti, Lorenzo, et al.
Published: (2025)
by: Proietti, Lorenzo, et al.
Published: (2025)
Span-Level Machine Translation Meta-Evaluation
by: Perrella, Stefano, et al.
Published: (2026)
by: Perrella, Stefano, et al.
Published: (2026)
Interpretable Coreference Resolution Evaluation Using Explicit Semantics
by: Gatti, Bruno, et al.
Published: (2026)
by: Gatti, Bruno, et al.
Published: (2026)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
by: Bonomo, Tommaso, et al.
Published: (2025)
by: Bonomo, Tommaso, et al.
Published: (2025)
Genie: Achieving Human Parity in Content-Grounded Datasets Generation
by: Yehudai, Asaf, et al.
Published: (2024)
by: Yehudai, Asaf, et al.
Published: (2024)
On the Evaluation Practices in Multilingual NLP: Can Machine Translation Offer an Alternative to Human Translations?
by: Choenni, Rochelle, et al.
Published: (2024)
by: Choenni, Rochelle, et al.
Published: (2024)
Maverick: Efficient and Accurate Coreference Resolution Defying Recent Trends
by: Martinelli, Giuliano, et al.
Published: (2024)
by: Martinelli, Giuliano, et al.
Published: (2024)
An Interdisciplinary Approach to Human-Centered Machine Translation
by: Carpuat, Marine, et al.
Published: (2025)
by: Carpuat, Marine, et al.
Published: (2025)
SimulPL: Aligning Human Preferences in Simultaneous Machine Translation
by: Yu, Donglei, et al.
Published: (2025)
by: Yu, Donglei, et al.
Published: (2025)
AutoML-guided Fusion of Entity and LLM-based Representations for Document Classification
by: Koloski, Boshko, et al.
Published: (2024)
by: Koloski, Boshko, et al.
Published: (2024)
Convergences and Divergences between Automatic Assessment and Human Evaluation: Insights from Comparing ChatGPT-Generated Translation and Neural Machine Translation
by: Jiang, Zhaokun, et al.
Published: (2024)
by: Jiang, Zhaokun, et al.
Published: (2024)
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
by: Raunak, Vikas, et al.
Published: (2023)
by: Raunak, Vikas, et al.
Published: (2023)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
by: Zhang, Ran, et al.
Published: (2024)
by: Zhang, Ran, et al.
Published: (2024)
Enhancing Neural Machine Translation of Low-Resource Languages: Corpus Development, Human Evaluation and Explainable AI Architectures
by: Lankford, Séamus
Published: (2024)
by: Lankford, Séamus
Published: (2024)
Trainable Reference-Based Evaluation Metric for Identifying Quality of English-Gujarati Machine Translation System
by: Joshi, Nisheeth, et al.
Published: (2025)
by: Joshi, Nisheeth, et al.
Published: (2025)
BOOKCOREF: Coreference Resolution at Book Scale
by: Martinelli, Giuliano, et al.
Published: (2025)
by: Martinelli, Giuliano, et al.
Published: (2025)
Do Large Language Models Understand Word Senses?
by: Meconi, Domenico, et al.
Published: (2025)
by: Meconi, Domenico, et al.
Published: (2025)
Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward Model
by: He, Zhiwei, et al.
Published: (2024)
by: He, Zhiwei, et al.
Published: (2024)
ReLiK: Retrieve and LinK, Fast and Accurate Entity Linking and Relation Extraction on an Academic Budget
by: Orlando, Riccardo, et al.
Published: (2024)
by: Orlando, Riccardo, et al.
Published: (2024)
PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models
by: Hu, Sihao, et al.
Published: (2024)
by: Hu, Sihao, et al.
Published: (2024)
HumBEL: A Human-in-the-Loop Approach for Evaluating Demographic Factors of Language Models in Human-Machine Conversations
by: Sicilia, Anthony, et al.
Published: (2023)
by: Sicilia, Anthony, et al.
Published: (2023)
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
by: Proietti, Lorenzo, et al.
Published: (2026)
by: Proietti, Lorenzo, et al.
Published: (2026)
What Has Been Lost with Synthetic Evaluation?
by: Gill, Alexander, et al.
Published: (2025)
by: Gill, Alexander, et al.
Published: (2025)
Word Sense Linking: Disambiguating Outside the Sandbox
by: Bejgu, Andrei Stefan, et al.
Published: (2024)
by: Bejgu, Andrei Stefan, et al.
Published: (2024)
Evaluation of Machine Translation Based on Semantic Dependencies and Keywords
by: Yuan, Kewei, et al.
Published: (2024)
by: Yuan, Kewei, et al.
Published: (2024)
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation
by: Lyu, Boxuan, et al.
Published: (2025)
by: Lyu, Boxuan, et al.
Published: (2025)
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
by: Yan, Jianhao, et al.
Published: (2024)
by: Yan, Jianhao, et al.
Published: (2024)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
by: Kreutzer, Julia, et al.
Published: (2025)
by: Kreutzer, Julia, et al.
Published: (2025)
Strategies of Code-switching in Human-Machine Dialogs
by: Geckt, Dean, et al.
Published: (2025)
by: Geckt, Dean, et al.
Published: (2025)
EEG-to-Text Translation: A Model for Deciphering Human Brain Activity
by: Murad, Saydul Akbar, et al.
Published: (2025)
by: Murad, Saydul Akbar, et al.
Published: (2025)
Human-Instruction-Free LLM Self-Alignment with Limited Samples
by: Guo, Hongyi, et al.
Published: (2024)
by: Guo, Hongyi, et al.
Published: (2024)
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models
by: Elangovan, Aparna, et al.
Published: (2024)
by: Elangovan, Aparna, et al.
Published: (2024)
FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity
by: Jourdan, Fanny, et al.
Published: (2025)
by: Jourdan, Fanny, et al.
Published: (2025)
Efficient Machine Translation Corpus Generation: Integrating Human-in-the-Loop Post-Editing with Large Language Models
by: Yuksel, Kamer Ali, et al.
Published: (2025)
by: Yuksel, Kamer Ali, et al.
Published: (2025)
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
by: Chen, Zixun, et al.
Published: (2025)
by: Chen, Zixun, et al.
Published: (2025)
Toward Human-Centered Readability Evaluation
by: İlgen, Bahar, et al.
Published: (2025)
by: İlgen, Bahar, et al.
Published: (2025)
Similar Items
-
Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
by: Perrella, Stefano, et al.
Published: (2024) -
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics
by: Perrella, Stefano, et al.
Published: (2024) -
Estimating Machine Translation Difficulty
by: Proietti, Lorenzo, et al.
Published: (2025) -
Span-Level Machine Translation Meta-Evaluation
by: Perrella, Stefano, et al.
Published: (2026) -
Interpretable Coreference Resolution Evaluation Using Explicit Semantics
by: Gatti, Bruno, et al.
Published: (2026)