MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
Fuente:
arXiv
Saved in:
| Main Authors: | Juraska, Juraj, Deutsch, Daniel, Finkelstein, Mara, Freitag, Markus |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025)
by: Juraska, Juraj, et al.
Published: (2025)
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects
by: Deutsch, Daniel, et al.
Published: (2025)
by: Deutsch, Daniel, et al.
Published: (2025)
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025)
by: Riley, Parker, et al.
Published: (2025)
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
LLMRefine: Pinpointing and Refining Large Language Models via Fine-Grained Actionable Feedback
by: Xu, Wenda, et al.
Published: (2023)
by: Xu, Wenda, et al.
Published: (2023)
Mitigating Metric Bias in Minimum Bayes Risk Decoding
by: Kovacs, Geza, et al.
Published: (2024)
by: Kovacs, Geza, et al.
Published: (2024)
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
by: Bhaskar, Yash, et al.
Published: (2025)
by: Bhaskar, Yash, et al.
Published: (2025)
From SALAMANDRA to SALAMANDRATA: BSC Submission for WMT25 General Machine Translation Shared Task
by: Gilabert, Javier Garcia, et al.
Published: (2025)
by: Gilabert, Javier Garcia, et al.
Published: (2025)
Findings of the WMT 2024 Shared Task on Chat Translation
by: Mohammed, Wafaa, et al.
Published: (2024)
by: Mohammed, Wafaa, et al.
Published: (2024)
JGU Mainz's Submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA
by: Saadi, Hossain Shaikh, et al.
Published: (2025)
by: Saadi, Hossain Shaikh, et al.
Published: (2025)
Findings of the WMT 2024 Shared Task on Discourse-Level Literary Translation
by: Wang, Longyue, et al.
Published: (2024)
by: Wang, Longyue, et al.
Published: (2024)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
Cogs in a Machine, Doing What They're Meant to Do -- The AMI Submission to the WMT24 General Translation Task
by: Jasonarson, Atli, et al.
Published: (2024)
by: Jasonarson, Atli, et al.
Published: (2024)
Preliminary WMT24 Ranking of General MT Systems and LLMs
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms
by: Trabelsi, Firas, et al.
Published: (2024)
by: Trabelsi, Firas, et al.
Published: (2024)
NLIP_Lab-IITH Low-Resource MT System for WMT24 Indic MT Shared Task
by: Sahoo, Pramit, et al.
Published: (2024)
by: Sahoo, Pramit, et al.
Published: (2024)
Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization
by: Li, Zilong
Published: (2024)
by: Li, Zilong
Published: (2024)
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025)
by: Shayegh, Behzad, et al.
Published: (2025)
MBR and QE Finetuning: Training-time Distillation of the Best and Most Expensive Decoding Methods
by: Finkelstein, Mara, et al.
Published: (2023)
by: Finkelstein, Mara, et al.
Published: (2023)
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
IKUN for WMT24 General MT Task: LLMs Are here for Multilingual Machine Translation
by: Liao, Baohao, et al.
Published: (2024)
by: Liao, Baohao, et al.
Published: (2024)
Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics
by: Zhang, Yilin, et al.
Published: (2025)
by: Zhang, Yilin, et al.
Published: (2025)
In2x at WMT25 Translation Task
by: Pang, Lei, et al.
Published: (2025)
by: Pang, Lei, et al.
Published: (2025)
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
by: Thompson, Brian, et al.
Published: (2024)
by: Thompson, Brian, et al.
Published: (2024)
TranslateGemma Technical Report
by: Finkelstein, Mara, et al.
Published: (2026)
by: Finkelstein, Mara, et al.
Published: (2026)
Sheffield's Submission to the AmericasNLP Shared Task on Machine Translation into Indigenous Languages
by: Gow-Smith, Edward, et al.
Published: (2023)
by: Gow-Smith, Edward, et al.
Published: (2023)
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
by: Xu, Wenda, et al.
Published: (2025)
by: Xu, Wenda, et al.
Published: (2025)
Preliminary Ranking of WMT25 General Machine Translation Systems
by: Kocmi, Tom, et al.
Published: (2025)
by: Kocmi, Tom, et al.
Published: (2025)
WMT24 Test Suite: Gender Resolution in Speaker-Listener Dialogue Roles
by: Dawkins, Hillary, et al.
Published: (2024)
by: Dawkins, Hillary, et al.
Published: (2024)
Learning from others' mistakes: Finetuning machine translation models with span-level error annotations
by: Zhang, Lily H., et al.
Published: (2024)
by: Zhang, Lily H., et al.
Published: (2024)
Finding Replicable Human Evaluations via Stable Ranking Probability
by: Riley, Parker, et al.
Published: (2024)
by: Riley, Parker, et al.
Published: (2024)
HW-TSC's Submission to the CCMT 2024 Machine Translation Tasks
by: Wu, Zhanglin, et al.
Published: (2024)
by: Wu, Zhanglin, et al.
Published: (2024)
Expanding the WMT24++ Benchmark with Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, and Vallader
by: Vamvas, Jannis, et al.
Published: (2025)
by: Vamvas, Jannis, et al.
Published: (2025)
Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
by: Tomani, Christian, et al.
Published: (2023)
by: Tomani, Christian, et al.
Published: (2023)
Exploring Parameter-Efficient Fine-Tuning and Backtranslation for the WMT 25 General Translation Task
by: Fujita, Felipe, et al.
Published: (2025)
by: Fujita, Felipe, et al.
Published: (2025)
Culturally Grounded Physical Commonsense Reasoning in Italian and English: A Submission to the MRL 2025 Shared Task
by: De Santis, Marco, et al.
Published: (2025)
by: De Santis, Marco, et al.
Published: (2025)
Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data
by: Liu, Zhongtao, et al.
Published: (2024)
by: Liu, Zhongtao, et al.
Published: (2024)
Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
by: Kocyigit, Muhammed Yusuf, et al.
Published: (2025)
MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
by: Winata, Genta Indra, et al.
Published: (2024)
by: Winata, Genta Indra, et al.
Published: (2024)
Similar Items
-
MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
by: Juraska, Juraj, et al.
Published: (2025) -
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects
by: Deutsch, Daniel, et al.
Published: (2025) -
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
by: Finkelstein, Mara, et al.
Published: (2024) -
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025) -
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025)