M-Prometheus: A Suite of Open Multilingual LLM Judges

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pombal, José, Yoon, Dongkeun, Fernandes, Patrick, Wu, Ian, Kim, Seungone, Rei, Ricardo, Neubig, Graham, Martins, André F. T.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912677416140800
author Pombal, José
Yoon, Dongkeun
Fernandes, Patrick
Wu, Ian
Kim, Seungone
Rei, Ricardo
Neubig, Graham
Martins, André F. T.
author_facet Pombal, José
Yoon, Dongkeun
Fernandes, Patrick
Wu, Ian
Kim, Seungone
Rei, Ricardo
Neubig, Graham
Martins, André F. T.
contents The use of language models for automatically evaluating long-form text (LLM-as-a-judge) is becoming increasingly common, yet most LLM judges are optimized exclusively for English, with strategies for enhancing their multilingual evaluation capabilities remaining largely unexplored in the current literature. This has created a disparity in the quality of automatic evaluation methods for non-English languages, ultimately hindering the development of models with better multilingual capabilities. To bridge this gap, we introduce M-Prometheus, a suite of open-weight LLM judges ranging from 3B to 14B parameters that can provide both direct assessment and pairwise comparison feedback on multilingual outputs. M-Prometheus models outperform state-of-the-art open LLM judges on multilingual reward benchmarks spanning more than 20 languages, as well as on literary machine translation (MT) evaluation covering 4 language pairs. Furthermore, M-Prometheus models can be leveraged at decoding time to significantly improve generated outputs across all 3 tested languages, showcasing their utility for the development of better multilingual models. Lastly, through extensive ablations, we identify the key factors for obtaining an effective multilingual judge, including backbone model selection and training on synthetic multilingual feedback data instead of translated data. We release our models, training dataset, and code.
format Preprint
id arxiv_https___arxiv_org_abs_2504_04953
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle M-Prometheus: A Suite of Open Multilingual LLM Judges
Pombal, José
Yoon, Dongkeun
Fernandes, Patrick
Wu, Ian
Kim, Seungone
Rei, Ricardo
Neubig, Graham
Martins, André F. T.
Computation and Language
Artificial Intelligence
The use of language models for automatically evaluating long-form text (LLM-as-a-judge) is becoming increasingly common, yet most LLM judges are optimized exclusively for English, with strategies for enhancing their multilingual evaluation capabilities remaining largely unexplored in the current literature. This has created a disparity in the quality of automatic evaluation methods for non-English languages, ultimately hindering the development of models with better multilingual capabilities. To bridge this gap, we introduce M-Prometheus, a suite of open-weight LLM judges ranging from 3B to 14B parameters that can provide both direct assessment and pairwise comparison feedback on multilingual outputs. M-Prometheus models outperform state-of-the-art open LLM judges on multilingual reward benchmarks spanning more than 20 languages, as well as on literary machine translation (MT) evaluation covering 4 language pairs. Furthermore, M-Prometheus models can be leveraged at decoding time to significantly improve generated outputs across all 3 tested languages, showcasing their utility for the development of better multilingual models. Lastly, through extensive ablations, we identify the key factors for obtaining an effective multilingual judge, including backbone model selection and training on synthetic multilingual feedback data instead of translated data. We release our models, training dataset, and code.
title M-Prometheus: A Suite of Open Multilingual LLM Judges
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.04953