M-Prometheus: A Suite of Open Multilingual LLM Judges
Fuente:
arXiv
Saved in:
| Main Authors: | Pombal, José, Yoon, Dongkeun, Fernandes, Patrick, Wu, Ian, Kim, Seungone, Rei, Ricardo, Neubig, Graham, Martins, André F. T. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LangBridge: Multilingual Reasoning Without Multilingual Supervision
by: Yoon, Dongkeun, et al.
Published: (2024)
by: Yoon, Dongkeun, et al.
Published: (2024)
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
by: Pombal, José, et al.
Published: (2026)
by: Pombal, José, et al.
Published: (2026)
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
by: Kim, Seungone, et al.
Published: (2024)
by: Kim, Seungone, et al.
Published: (2024)
Better Instruction-Following Through Minimum Bayes Risk
by: Wu, Ian, et al.
Published: (2024)
by: Wu, Ian, et al.
Published: (2024)
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
by: Lee, Seongyun, et al.
Published: (2024)
by: Lee, Seongyun, et al.
Published: (2024)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
Adding Chocolate to Mint: Mitigating Metric Interference in Machine Translation
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
by: Rei, Ricardo, et al.
Published: (2025)
by: Rei, Ricardo, et al.
Published: (2025)
xTower: A Multilingual LLM for Explaining and Correcting Translation Errors
by: Treviso, Marcos, et al.
Published: (2024)
by: Treviso, Marcos, et al.
Published: (2024)
Tower: An Open Multilingual Large Language Model for Translation-Related Tasks
by: Alves, Duarte M., et al.
Published: (2024)
by: Alves, Duarte M., et al.
Published: (2024)
EuroLLM: Multilingual Language Models for Europe
by: Martins, Pedro Henrique, et al.
Published: (2024)
by: Martins, Pedro Henrique, et al.
Published: (2024)
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
by: Yue, Xiang, et al.
Published: (2024)
by: Yue, Xiang, et al.
Published: (2024)
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
by: Sutawika, Lintang, et al.
Published: (2026)
by: Sutawika, Lintang, et al.
Published: (2026)
A Context-aware Framework for Translation-mediated Conversations
by: Pombal, José, et al.
Published: (2024)
by: Pombal, José, et al.
Published: (2024)
Do LLMs Understand Your Translations? Evaluating Paragraph-level MT with Question Answering
by: Fernandes, Patrick, et al.
Published: (2025)
by: Fernandes, Patrick, et al.
Published: (2025)
Is Context Helpful for Chat Translation Evaluation?
by: Agrawal, Sweta, et al.
Published: (2024)
by: Agrawal, Sweta, et al.
Published: (2024)
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
by: Kim, Seungone, et al.
Published: (2023)
by: Kim, Seungone, et al.
Published: (2023)
Reasoning Models Better Express Their Confidence
by: Yoon, Dongkeun, et al.
Published: (2025)
by: Yoon, Dongkeun, et al.
Published: (2025)
What Is Missing in Multilingual Visual Reasoning and How to Fix It
by: Song, Yueqi, et al.
Published: (2024)
by: Song, Yueqi, et al.
Published: (2024)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
by: Huan, Maggie, et al.
Published: (2025)
by: Huan, Maggie, et al.
Published: (2025)
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation
by: Iyer, Vivek, et al.
Published: (2025)
by: Iyer, Vivek, et al.
Published: (2025)
Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
by: Ramos, Miguel Moura, et al.
Published: (2025)
by: Ramos, Miguel Moura, et al.
Published: (2025)
MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
Scaling Evaluation-time Compute with Reasoning Models as Evaluators
by: Kim, Seungone, et al.
Published: (2025)
by: Kim, Seungone, et al.
Published: (2025)
SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following
by: Canaverde, Beatriz, et al.
Published: (2026)
by: Canaverde, Beatriz, et al.
Published: (2026)
Can Language Models Evaluate Human Written Text? Case Study on Korean Student Writing for Education
by: Kim, Seungyoon, et al.
Published: (2024)
by: Kim, Seungyoon, et al.
Published: (2024)
EuroLLM-9B: Technical Report
by: Martins, Pedro Henrique, et al.
Published: (2025)
by: Martins, Pedro Henrique, et al.
Published: (2025)
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
How Reliable is Multilingual LLM-as-a-Judge?
by: Fu, Xiyan, et al.
Published: (2025)
by: Fu, Xiyan, et al.
Published: (2025)
Checklist Engineering Empowers Multilingual LLM Judges
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2025)
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2025)
Can Automatic Metrics Assess High-Quality Translations?
by: Agrawal, Sweta, et al.
Published: (2024)
by: Agrawal, Sweta, et al.
Published: (2024)
EuroLLM-22B: Technical Report
by: Ramos, Miguel Moura, et al.
Published: (2026)
by: Ramos, Miguel Moura, et al.
Published: (2026)
Go-Browse: Training Web Agents with Structured Exploration
by: Gandhi, Apurva, et al.
Published: (2025)
by: Gandhi, Apurva, et al.
Published: (2025)
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models
by: Tjuatja, Lindia, et al.
Published: (2025)
by: Tjuatja, Lindia, et al.
Published: (2025)
ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
by: Xu, Yiming, et al.
Published: (2025)
by: Xu, Yiming, et al.
Published: (2025)
Evaluating Language Models as Synthetic Data Generators
by: Kim, Seungone, et al.
Published: (2024)
by: Kim, Seungone, et al.
Published: (2024)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2024)
by: Kim, Eunsu, et al.
Published: (2024)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
by: Freitas, Miguel Monte e, et al.
Published: (2026)
by: Freitas, Miguel Monte e, et al.
Published: (2026)
Similar Items
-
LangBridge: Multilingual Reasoning Without Multilingual Supervision
by: Yoon, Dongkeun, et al.
Published: (2024) -
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
by: Pombal, José, et al.
Published: (2026) -
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
by: Kim, Seungone, et al.
Published: (2024) -
Better Instruction-Following Through Minimum Bayes Risk
by: Wu, Ian, et al.
Published: (2024) -
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
by: Lee, Seongyun, et al.
Published: (2024)