MetricX-25 and GemSpanEval: Google Translate Submissions to the WMT25 Evaluation Shared Task
Fuente:
arXiv
Saved in:
| Main Authors: | Juraska, Juraj, Domhan, Tobias, Finkelstein, Mara, Nakagawa, Tetsuji, Kovacs, Geza, Deutsch, Daniel, Wang, Pidong, Freitag, Markus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
by: Juraska, Juraj, et al.
Published: (2024)
by: Juraska, Juraj, et al.
Published: (2024)
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects
by: Deutsch, Daniel, et al.
Published: (2025)
by: Deutsch, Daniel, et al.
Published: (2025)
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025)
by: Riley, Parker, et al.
Published: (2025)
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025)
by: Zouhar, Vilém, et al.
Published: (2025)
Mitigating Metric Bias in Minimum Bayes Risk Decoding
by: Kovacs, Geza, et al.
Published: (2024)
by: Kovacs, Geza, et al.
Published: (2024)
Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation
by: Shayegh, Behzad, et al.
Published: (2025)
by: Shayegh, Behzad, et al.
Published: (2025)
From SALAMANDRA to SALAMANDRATA: BSC Submission for WMT25 General Machine Translation Shared Task
by: Gilabert, Javier Garcia, et al.
Published: (2025)
by: Gilabert, Javier Garcia, et al.
Published: (2025)
TranslateGemma Technical Report
by: Finkelstein, Mara, et al.
Published: (2026)
by: Finkelstein, Mara, et al.
Published: (2026)
LLMRefine: Pinpointing and Refining Large Language Models via Fine-Grained Actionable Feedback
by: Xu, Wenda, et al.
Published: (2023)
by: Xu, Wenda, et al.
Published: (2023)
JGU Mainz's Submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA
by: Saadi, Hossain Shaikh, et al.
Published: (2025)
by: Saadi, Hossain Shaikh, et al.
Published: (2025)
In2x at WMT25 Translation Task
by: Pang, Lei, et al.
Published: (2025)
by: Pang, Lei, et al.
Published: (2025)
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
by: Bhaskar, Yash, et al.
Published: (2025)
by: Bhaskar, Yash, et al.
Published: (2025)
Preliminary Ranking of WMT25 General Machine Translation Systems
by: Kocmi, Tom, et al.
Published: (2025)
by: Kocmi, Tom, et al.
Published: (2025)
Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics
by: Zhang, Yilin, et al.
Published: (2025)
by: Zhang, Yilin, et al.
Published: (2025)
Mind the Gap... or Not? How Translation Errors and Evaluation Details Skew Multilingual Results
by: Peter, Jan-Thorsten, et al.
Published: (2025)
by: Peter, Jan-Thorsten, et al.
Published: (2025)
Findings of the WMT 2024 Shared Task on Chat Translation
by: Mohammed, Wafaa, et al.
Published: (2024)
by: Mohammed, Wafaa, et al.
Published: (2024)
Exploring Parameter-Efficient Fine-Tuning and Backtranslation for the WMT 25 General Translation Task
by: Fujita, Felipe, et al.
Published: (2025)
by: Fujita, Felipe, et al.
Published: (2025)
Choose the Final Translation from NMT and LLM hypotheses Using MBR Decoding: HW-TSC's Submission to the WMT24 General MT Shared Task
by: Wu, Zhanglin, et al.
Published: (2024)
by: Wu, Zhanglin, et al.
Published: (2024)
Findings of the WMT 2024 Shared Task on Discourse-Level Literary Translation
by: Wang, Longyue, et al.
Published: (2024)
by: Wang, Longyue, et al.
Published: (2024)
Pulsation-driven helium transport as a potential source of the Blazhko effect
by: Kovacs, Geza
Published: (2026)
by: Kovacs, Geza
Published: (2026)
Digging Deeper for RR Lyrae Stars with Low Modulation Amplitudes
by: Kovacs, Geza
Published: (2025)
by: Kovacs, Geza
Published: (2025)
Same evaluation, more tokens: On the effect of input length for machine translation evaluation using Large Language Models
by: Domhan, Tobias, et al.
Published: (2025)
by: Domhan, Tobias, et al.
Published: (2025)
Cogs in a Machine, Doing What They're Meant to Do -- The AMI Submission to the WMT24 General Translation Task
by: Jasonarson, Atli, et al.
Published: (2024)
by: Jasonarson, Atli, et al.
Published: (2024)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
Enhancing Human Evaluation in Machine Translation with Comparative Judgment
by: Song, Yixiao, et al.
Published: (2025)
by: Song, Yixiao, et al.
Published: (2025)
Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms
by: Trabelsi, Firas, et al.
Published: (2024)
by: Trabelsi, Firas, et al.
Published: (2024)
A Shocking Amount of the Web is Machine Translated: Insights from Multi-Way Parallelism
by: Thompson, Brian, et al.
Published: (2024)
by: Thompson, Brian, et al.
Published: (2024)
Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
by: Dadkhahi, Hamid, et al.
Published: (2025)
by: Dadkhahi, Hamid, et al.
Published: (2025)
Evaluating the efficacy of combined flap coverage, antibiotic‐loaded bone cement and negative pressure irrigation in traumatic osteomyelitis management
by: Pidong Liu, et al.
Published: (2024)
by: Pidong Liu, et al.
Published: (2024)
Sheffield's Submission to the AmericasNLP Shared Task on Machine Translation into Indigenous Languages
by: Gow-Smith, Edward, et al.
Published: (2023)
by: Gow-Smith, Edward, et al.
Published: (2023)
AIxcellent Vibes at GermEval 2025 Shared Task on Candy Speech Detection: Improving Model Performance by Span-Level Training
by: Thelen, Christian Rene, et al.
Published: (2025)
by: Thelen, Christian Rene, et al.
Published: (2025)
Interlocal Adaptations to Climate Change in East and Southeast Asia Sharing Lessons of Agriculture, Disaster Risk Reduction, and Resource Management
by: Tetsuji Ito,
by: Tetsuji Ito,
Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model
by: Tomani, Christian, et al.
Published: (2023)
by: Tomani, Christian, et al.
Published: (2023)
The last stage of development: The restructuring and plasticity of the cortex during adolescence especially at puberty
by: Janice M. Juraska
Published: (2024)
by: Janice M. Juraska
Published: (2024)
MBR and QE Finetuning: Training-time Distillation of the Best and Most Expensive Decoding Methods
by: Finkelstein, Mara, et al.
Published: (2023)
by: Finkelstein, Mara, et al.
Published: (2023)
Secondary eclipses of two brown dwarfs in the K2 fields: detection by multiple dataset merging
by: Kovacs, Geza, et al.
Published: (2025)
by: Kovacs, Geza, et al.
Published: (2025)
Pretraining Strategies using Monolingual and Parallel Data for Low-Resource Machine Translation
by: Nguefack, Idriss Nguepi, et al.
Published: (2025)
by: Nguefack, Idriss Nguepi, et al.
Published: (2025)
IKUN for WMT24 General MT Task: LLMs Are here for Multilingual Machine Translation
by: Liao, Baohao, et al.
Published: (2024)
by: Liao, Baohao, et al.
Published: (2024)
Preliminary WMT24 Ranking of General MT Systems and LLMs
by: Kocmi, Tom, et al.
Published: (2024)
by: Kocmi, Tom, et al.
Published: (2024)
Similar Items
-
MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task
by: Juraska, Juraj, et al.
Published: (2024) -
From Jack of All Trades to Master of One: Specializing LLM-based Autoraters to a Test Set
by: Finkelstein, Mara, et al.
Published: (2024) -
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects
by: Deutsch, Daniel, et al.
Published: (2025) -
MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
by: Riley, Parker, et al.
Published: (2025) -
Generating Difficult-to-Translate Texts
by: Zouhar, Vilém, et al.
Published: (2025)