MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
Fuente:
arXiv
Saved in:
| Main Authors: | Riley, Parker, Deutsch, Daniel, Finkelstein, Mara, DiIanni, Colten, Juraska, Juraj, Freitag, Markus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Don't Sweat the Small Stuff: Segment-Level Meta-Evaluation Based on Pairwise Difference Correlation
by: DiIanni, Colten, et al.
Published: (2025)
by: DiIanni, Colten, et al.
Published: (2025)
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
by: Juraska, Juraj, et al.
Published: (2024)
by: Juraska, Juraj, et al.
Published: (2024)
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025)
by: Juraska, Juraj, et al.
Published: (2025)
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data
by: Liu, Zhongtao, et al.
Published: (2024)
by: Liu, Zhongtao, et al.
Published: (2024)
LLMRefine: Pinpointing and Refining Large Language Models via Fine-Grained Actionable Feedback
by: Xu, Wenda, et al.
Published: (2023)
by: Xu, Wenda, et al.
Published: (2023)
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025)
by: Shayegh, Behzad, et al.
Published: (2025)
Finding Replicable Human Evaluations via Stable Ranking Probability
by: Riley, Parker, et al.
Published: (2024)
by: Riley, Parker, et al.
Published: (2024)
TranslateGemma Technical Report
by: Finkelstein, Mara, et al.
Published: (2026)
by: Finkelstein, Mara, et al.
Published: (2026)
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects
by: Deutsch, Daniel, et al.
Published: (2025)
by: Deutsch, Daniel, et al.
Published: (2025)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators
by: Lu, Qingyu, et al.
Published: (2024)
by: Lu, Qingyu, et al.
Published: (2024)
MQM-Chat: Multidimensional Quality Metrics for Chat Translation
by: Li, Yunmeng, et al.
Published: (2024)
by: Li, Yunmeng, et al.
Published: (2024)
Mitigating Metric Bias in Minimum Bayes Risk Decoding
by: Kovacs, Geza, et al.
Published: (2024)
by: Kovacs, Geza, et al.
Published: (2024)
Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms
by: Trabelsi, Firas, et al.
Published: (2024)
by: Trabelsi, Firas, et al.
Published: (2024)
MAATS: A Multi-Agent Automated Translation System Based on MQM Evaluation
by: Wang, George, et al.
Published: (2025)
by: Wang, George, et al.
Published: (2025)
MBR and QE Finetuning: Training-time Distillation of the Best and Most Expensive Decoding Methods
by: Finkelstein, Mara, et al.
Published: (2023)
by: Finkelstein, Mara, et al.
Published: (2023)
Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
by: Tomani, Christian, et al.
Published: (2023)
by: Tomani, Christian, et al.
Published: (2023)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
Pretraining Strategies using Monolingual and Parallel Data for Low-Resource Machine Translation
by: Nguefack, Idriss Nguepi, et al.
Published: (2025)
by: Nguefack, Idriss Nguepi, et al.
Published: (2025)
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
The Multi-Range Theory of Translation Quality Measurement: MQM scoring models and Statistical Quality Control
by: Lommel, Arle, et al.
Published: (2024)
by: Lommel, Arle, et al.
Published: (2024)
Learning from others' mistakes: Finetuning machine translation models with span-level error annotations
by: Zhang, Lily H., et al.
Published: (2024)
by: Zhang, Lily H., et al.
Published: (2024)
Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts
by: Briakou, Eleftheria, et al.
Published: (2024)
by: Briakou, Eleftheria, et al.
Published: (2024)
Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
by: Peter, Jan-Thorsten, et al.
Published: (2025)
by: Peter, Jan-Thorsten, et al.
Published: (2025)
Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations
by: Tan, Shaomu, et al.
Published: (2025)
by: Tan, Shaomu, et al.
Published: (2025)
Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
by: Dadkhahi, Hamid, et al.
Published: (2025)
by: Dadkhahi, Hamid, et al.
Published: (2025)
Preliminary Ranking of WMT25 General Machine Translation Systems
by: Kocmi, Tom, et al.
Published: (2025)
by: Kocmi, Tom, et al.
Published: (2025)
ConGA: Guidelines for Contextual Gender Annotation. A Framework for Annotating Gender in Machine Translation
by: Rescigno, Argentina Anna, et al.
Published: (2026)
by: Rescigno, Argentina Anna, et al.
Published: (2026)
You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation
by: Flamich, Gergely, et al.
Published: (2025)
by: Flamich, Gergely, et al.
Published: (2025)
Large Language Models as Annotators for Machine Translation Quality Estimation
by: Wang, Sidi, et al.
Published: (2026)
by: Wang, Sidi, et al.
Published: (2026)
Finnish SQuAD: A Simple Approach to Machine Translation of Span Annotations
by: Nuutinen, Emil, et al.
Published: (2025)
by: Nuutinen, Emil, et al.
Published: (2025)
SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic Languages
by: Liu, Hannah, et al.
Published: (2025)
by: Liu, Hannah, et al.
Published: (2025)
Translation of Multifaceted Data without Re-Training of Machine Translation Systems
by: Moon, Hyeonseok, et al.
Published: (2024)
by: Moon, Hyeonseok, et al.
Published: (2024)
"Be My Cheese?": Cultural Nuance Benchmarking for Machine Translation in Multilingual LLMs
by: Van Doren, Madison, et al.
Published: (2026)
by: Van Doren, Madison, et al.
Published: (2026)
Hierarchical Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text
by: Morbiato, Filippo, et al.
Published: (2026)
by: Morbiato, Filippo, et al.
Published: (2026)
Similar Items
-
Don't Sweat the Small Stuff: Segment-Level Meta-Evaluation Based on Pairwise Difference Correlation
by: DiIanni, Colten, et al.
Published: (2025) -
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025) -
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
by: Juraska, Juraj, et al.
Published: (2024) -
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
by: Finkelstein, Mara, et al.
Published: (2024) -
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025)