How Reliable is Multilingual LLM-as-a-Judge?
Fuente:
arXiv
Salvato in:
| Autori principali: | Fu, Xiyan, Liu, Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization
di: Fu, Xiyan, et al.
Pubblicazione: (2026)
di: Fu, Xiyan, et al.
Pubblicazione: (2026)
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
di: Yamauchi, Yusuke, et al.
Pubblicazione: (2025)
di: Yamauchi, Yusuke, et al.
Pubblicazione: (2025)
Checklist Engineering Empowers Multilingual LLM Judges
di: Mohammadkhani, Mohammad Ghiasvand, et al.
Pubblicazione: (2025)
di: Mohammadkhani, Mohammad Ghiasvand, et al.
Pubblicazione: (2025)
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
di: Schroeder, Kayla, et al.
Pubblicazione: (2024)
di: Schroeder, Kayla, et al.
Pubblicazione: (2024)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
di: Wang, Yidong, et al.
Pubblicazione: (2025)
di: Wang, Yidong, et al.
Pubblicazione: (2025)
Exploring Continual Learning of Compositional Generalization in NLI
di: Fu, Xiyan, et al.
Pubblicazione: (2024)
di: Fu, Xiyan, et al.
Pubblicazione: (2024)
The Mystery of Compositional Generalization in Graph-based Generative Commonsense Reasoning
di: Fu, Xiyan, et al.
Pubblicazione: (2024)
di: Fu, Xiyan, et al.
Pubblicazione: (2024)
M-Prometheus: A Suite of Open Multilingual LLM Judges
di: Pombal, José, et al.
Pubblicazione: (2025)
di: Pombal, José, et al.
Pubblicazione: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
di: Bologna, Federica, et al.
Pubblicazione: (2026)
di: Bologna, Federica, et al.
Pubblicazione: (2026)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
di: Zhang, Hongbin, et al.
Pubblicazione: (2026)
di: Zhang, Hongbin, et al.
Pubblicazione: (2026)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
di: Hong, Yihan, et al.
Pubblicazione: (2026)
di: Hong, Yihan, et al.
Pubblicazione: (2026)
How to Correctly Report LLM-as-a-Judge Evaluations
di: Lee, Chungpa, et al.
Pubblicazione: (2025)
di: Lee, Chungpa, et al.
Pubblicazione: (2025)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
di: Alam, Firoj, et al.
Pubblicazione: (2026)
di: Alam, Firoj, et al.
Pubblicazione: (2026)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
di: Zhou, Yuhang, et al.
Pubblicazione: (2025)
di: Zhou, Yuhang, et al.
Pubblicazione: (2025)
DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection
di: Wu, Junchao, et al.
Pubblicazione: (2026)
di: Wu, Junchao, et al.
Pubblicazione: (2026)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
di: Yang, Bo, et al.
Pubblicazione: (2026)
di: Yang, Bo, et al.
Pubblicazione: (2026)
Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
di: Zhang, Xingjian, et al.
Pubblicazione: (2025)
di: Zhang, Xingjian, et al.
Pubblicazione: (2025)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
di: Tang, Zhenwei, et al.
Pubblicazione: (2026)
di: Tang, Zhenwei, et al.
Pubblicazione: (2026)
EnsemJudge: Enhancing Reliability in Chinese LLM-Generated Text Detection through Diverse Model Ensembles
di: Wang, Zhuoshang, et al.
Pubblicazione: (2026)
di: Wang, Zhuoshang, et al.
Pubblicazione: (2026)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
di: Gupta, Manan, et al.
Pubblicazione: (2026)
di: Gupta, Manan, et al.
Pubblicazione: (2026)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
di: Belmadani, Ikram, et al.
Pubblicazione: (2026)
di: Belmadani, Ikram, et al.
Pubblicazione: (2026)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
di: Huang, Hui, et al.
Pubblicazione: (2024)
di: Huang, Hui, et al.
Pubblicazione: (2024)
A Survey on LLM-as-a-Judge
di: Gu, Jiawei, et al.
Pubblicazione: (2024)
di: Gu, Jiawei, et al.
Pubblicazione: (2024)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
di: Schwinn, Leo, et al.
Pubblicazione: (2026)
di: Schwinn, Leo, et al.
Pubblicazione: (2026)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
di: Wei, Hui, et al.
Pubblicazione: (2024)
di: Wei, Hui, et al.
Pubblicazione: (2024)
What to Format and How: A Benchmark and Workflow Approach for Document Formatting
di: Rao, Shihao, et al.
Pubblicazione: (2026)
di: Rao, Shihao, et al.
Pubblicazione: (2026)
Think-J: Learning to Think for Generative LLM-as-a-Judge
di: Huang, Hui, et al.
Pubblicazione: (2025)
di: Huang, Hui, et al.
Pubblicazione: (2025)
Can LLM be a Personalized Judge?
di: Dong, Yijiang River, et al.
Pubblicazione: (2024)
di: Dong, Yijiang River, et al.
Pubblicazione: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
di: Shi, Lin, et al.
Pubblicazione: (2024)
di: Shi, Lin, et al.
Pubblicazione: (2024)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
di: Bellibatlu, Rohith Reddy, et al.
Pubblicazione: (2026)
di: Bellibatlu, Rohith Reddy, et al.
Pubblicazione: (2026)
Facts are Harder Than Opinions -- A Multilingual, Comparative Analysis of LLM-Based Fact-Checking Reliability
di: Saju, Lorraine, et al.
Pubblicazione: (2025)
di: Saju, Lorraine, et al.
Pubblicazione: (2025)
Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
di: Wei, Zhipeng, et al.
Pubblicazione: (2024)
di: Wei, Zhipeng, et al.
Pubblicazione: (2024)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
di: Zhu, Ziyi, et al.
Pubblicazione: (2026)
di: Zhu, Ziyi, et al.
Pubblicazione: (2026)
Evaluating Scoring Bias in LLM-as-a-Judge
di: Li, Qingquan, et al.
Pubblicazione: (2025)
di: Li, Qingquan, et al.
Pubblicazione: (2025)
The Necessity of Setting Temperature in LLM-as-a-Judge
di: Li, Lujun, et al.
Pubblicazione: (2026)
di: Li, Lujun, et al.
Pubblicazione: (2026)
Self-Preference Bias in LLM-as-a-Judge
di: Wataoka, Koki, et al.
Pubblicazione: (2024)
di: Wataoka, Koki, et al.
Pubblicazione: (2024)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
di: Chen, Junjie, et al.
Pubblicazione: (2026)
di: Chen, Junjie, et al.
Pubblicazione: (2026)
One Token to Fool LLM-as-a-Judge
di: Zhao, Yulai, et al.
Pubblicazione: (2025)
di: Zhao, Yulai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization
di: Fu, Xiyan, et al.
Pubblicazione: (2026) -
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
di: Yamauchi, Yusuke, et al.
Pubblicazione: (2025) -
Checklist Engineering Empowers Multilingual LLM Judges
di: Mohammadkhani, Mohammad Ghiasvand, et al.
Pubblicazione: (2025) -
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
di: Schroeder, Kayla, et al.
Pubblicazione: (2024) -
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
di: Wang, Yidong, et al.
Pubblicazione: (2025)