Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!
Fuente:
arXiv
Saved in:
| Main Authors: | Perrella, Stefano, Proietti, Lorenzo, Scirè, Alessandro, Barba, Edoardo, Navigli, Roberto |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics
by: Perrella, Stefano, et al.
Published: (2024)
by: Perrella, Stefano, et al.
Published: (2024)
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress
by: Proietti, Lorenzo, et al.
Published: (2025)
by: Proietti, Lorenzo, et al.
Published: (2025)
Estimating Machine Translation Difficulty
by: Proietti, Lorenzo, et al.
Published: (2025)
by: Proietti, Lorenzo, et al.
Published: (2025)
Span-Level Machine Translation Meta-Evaluation
by: Perrella, Stefano, et al.
Published: (2026)
by: Perrella, Stefano, et al.
Published: (2026)
Maverick: Efficient and Accurate Coreference Resolution Defying Recent Trends
by: Martinelli, Giuliano, et al.
Published: (2024)
by: Martinelli, Giuliano, et al.
Published: (2024)
ReLiK: Retrieve and LinK, Fast and Accurate Entity Linking and Relation Extraction on an Academic Budget
by: Orlando, Riccardo, et al.
Published: (2024)
by: Orlando, Riccardo, et al.
Published: (2024)
FENICE: Factuality Evaluation of summarization based on Natural language Inference and Claim Extraction
by: Scirè, Alessandro, et al.
Published: (2024)
by: Scirè, Alessandro, et al.
Published: (2024)
Word Sense Linking: Disambiguating Outside the Sandbox
by: Bejgu, Andrei Stefan, et al.
Published: (2024)
by: Bejgu, Andrei Stefan, et al.
Published: (2024)
Interpretable Coreference Resolution Evaluation Using Explicit Semantics
by: Gatti, Bruno, et al.
Published: (2026)
by: Gatti, Bruno, et al.
Published: (2026)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
by: Bonomo, Tommaso, et al.
Published: (2025)
by: Bonomo, Tommaso, et al.
Published: (2025)
MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration
by: Anugraha, David, et al.
Published: (2024)
by: Anugraha, David, et al.
Published: (2024)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
by: Polák, Peter, et al.
Published: (2025)
by: Polák, Peter, et al.
Published: (2025)
AutoML-guided Fusion of Entity and LLM-based Representations for Document Classification
by: Koloski, Boshko, et al.
Published: (2024)
by: Koloski, Boshko, et al.
Published: (2024)
Trainable Reference-Based Evaluation Metric for Identifying Quality of English-Gujarati Machine Translation System
by: Joshi, Nisheeth, et al.
Published: (2025)
by: Joshi, Nisheeth, et al.
Published: (2025)
Truth or Mirage? Towards End-to-End Factuality Evaluation with LLM-Oasis
by: Scirè, Alessandro, et al.
Published: (2024)
by: Scirè, Alessandro, et al.
Published: (2024)
Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains
by: Zouhar, Vilém, et al.
Published: (2024)
by: Zouhar, Vilém, et al.
Published: (2024)
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
LQM: Linguistically Motivated Multidimensional Quality Metrics for Machine Translation
by: Magdy, Samar M., et al.
Published: (2026)
by: Magdy, Samar M., et al.
Published: (2026)
Right Answer, Wrong Score: Uncovering the Inconsistencies of LLM Evaluation in Multiple-Choice Question Answering
by: Molfese, Francesco Maria, et al.
Published: (2025)
by: Molfese, Francesco Maria, et al.
Published: (2025)
BOOKCOREF: Coreference Resolution at Book Scale
by: Martinelli, Giuliano, et al.
Published: (2025)
by: Martinelli, Giuliano, et al.
Published: (2025)
Do Large Language Models Understand Word Senses?
by: Meconi, Domenico, et al.
Published: (2025)
by: Meconi, Domenico, et al.
Published: (2025)
Textual Similarity as a Key Metric in Machine Translation Quality Estimation
by: Sun, Kun, et al.
Published: (2024)
by: Sun, Kun, et al.
Published: (2024)
Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation
by: Sun, Yirong, et al.
Published: (2024)
by: Sun, Yirong, et al.
Published: (2024)
An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation
by: Rusli, Andre, et al.
Published: (2024)
by: Rusli, Andre, et al.
Published: (2024)
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
by: Proietti, Lorenzo, et al.
Published: (2026)
by: Proietti, Lorenzo, et al.
Published: (2026)
Automatic Evaluation Metrics for Document-level Translation: Overview, Challenges and Trends
by: GUO, Jiaxin, et al.
Published: (2025)
by: GUO, Jiaxin, et al.
Published: (2025)
How to Evaluate Speech Translation with Source-Aware Neural MT Metrics
by: Cettolo, Mauro, et al.
Published: (2025)
by: Cettolo, Mauro, et al.
Published: (2025)
Emergent Communication Pretraining for Few-Shot Machine Translation
by: Li, Yaoyiran, et al.
Published: (2020)
by: Li, Yaoyiran, et al.
Published: (2020)
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025)
by: Shayegh, Behzad, et al.
Published: (2025)
Evaluation of Machine Translation Based on Semantic Dependencies and Keywords
by: Yuan, Kewei, et al.
Published: (2024)
by: Yuan, Kewei, et al.
Published: (2024)
On the Evaluation Practices in Multilingual NLP: Can Machine Translation Offer an Alternative to Human Translations?
by: Choenni, Rochelle, et al.
Published: (2024)
by: Choenni, Rochelle, et al.
Published: (2024)
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation
by: Ki, Dayeon, et al.
Published: (2025)
by: Ki, Dayeon, et al.
Published: (2025)
QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems
by: Yang, Yiliu, et al.
Published: (2025)
by: Yang, Yiliu, et al.
Published: (2025)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
by: Kreutzer, Julia, et al.
Published: (2025)
by: Kreutzer, Julia, et al.
Published: (2025)
FairTranslate: An English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity
by: Jourdan, Fanny, et al.
Published: (2025)
by: Jourdan, Fanny, et al.
Published: (2025)
SLM as Guardian: Pioneering AI Safety with Small Language Models
by: Kwon, Ohjoon, et al.
Published: (2024)
by: Kwon, Ohjoon, et al.
Published: (2024)
POMP: Probability-driven Meta-graph Prompter for LLMs in Low-resource Unsupervised Neural Machine Translation
by: Pan, Shilong, et al.
Published: (2024)
by: Pan, Shilong, et al.
Published: (2024)
Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics
by: Park, Jin Hyun, et al.
Published: (2025)
by: Park, Jin Hyun, et al.
Published: (2025)
Meta-Evaluating Local LLMs: Rethinking Performance Metrics for Serious Games
by: Isaza-Giraldo, Andrés, et al.
Published: (2025)
by: Isaza-Giraldo, Andrés, et al.
Published: (2025)
M-MAD: Multidimensional Multi-Agent Debate for Advanced Machine Translation Evaluation
by: Feng, Zhaopeng, et al.
Published: (2024)
by: Feng, Zhaopeng, et al.
Published: (2024)
Similar Items
-
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics
by: Perrella, Stefano, et al.
Published: (2024) -
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress
by: Proietti, Lorenzo, et al.
Published: (2025) -
Estimating Machine Translation Difficulty
by: Proietti, Lorenzo, et al.
Published: (2025) -
Span-Level Machine Translation Meta-Evaluation
by: Perrella, Stefano, et al.
Published: (2026) -
Maverick: Efficient and Accurate Coreference Resolution Defying Recent Trends
by: Martinelli, Giuliano, et al.
Published: (2024)