Checklist Engineering Empowers Multilingual LLM Judges
Fuente:
arXiv
Saved in:
| Main Authors: | Mohammadkhani, Mohammad Ghiasvand, Beigy, Hamid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-Shot Learning and Key Points Are All You Need for Automated Fact-Checking
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2024)
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2024)
Gap-Filling Prompting Enhances Code-Assisted Mathematical Reasoning
by: Mohammadkhani, Mohammad Ghiasvand
Published: (2024)
by: Mohammadkhani, Mohammad Ghiasvand
Published: (2024)
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
E2TP: Element to Tuple Prompting Improves Aspect Sentiment Tuple Prediction
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2024)
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2024)
Consistency Training by Synthetic Question Generation for Conversational Question Answering
by: Hemati, Hamed Hematian, et al.
Published: (2024)
by: Hemati, Hamed Hematian, et al.
Published: (2024)
Multi-BERT: Leveraging Adapters and Prompt Tuning for Low-Resource Multi-Domain Adaptation
by: Azad, Parham Abed, et al.
Published: (2024)
by: Azad, Parham Abed, et al.
Published: (2024)
AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judge
by: Zhou, Karen, et al.
Published: (2026)
by: Zhou, Karen, et al.
Published: (2026)
FaBERT: Pre-training BERT on Persian Blogs
by: Masumi, Mostafa, et al.
Published: (2024)
by: Masumi, Mostafa, et al.
Published: (2024)
How Reliable is Multilingual LLM-as-a-Judge?
by: Fu, Xiyan, et al.
Published: (2025)
by: Fu, Xiyan, et al.
Published: (2025)
Could Thinking Multilingually Empower LLM Reasoning?
by: Gao, Changjiang, et al.
Published: (2025)
by: Gao, Changjiang, et al.
Published: (2025)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
M-Prometheus: A Suite of Open Multilingual LLM Judges
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
Analysis of Blood Report Images Using General Purpose Vision-Language Models
by: Bakhsheshi, Nadia, et al.
Published: (2025)
by: Bakhsheshi, Nadia, et al.
Published: (2025)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
by: Zhang, Hongbin, et al.
Published: (2026)
by: Zhang, Hongbin, et al.
Published: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
by: Marioriyad, Arash, et al.
Published: (2026)
by: Marioriyad, Arash, et al.
Published: (2026)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
by: Yang, Bo, et al.
Published: (2026)
by: Yang, Bo, et al.
Published: (2026)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
by: Belmadani, Ikram, et al.
Published: (2026)
by: Belmadani, Ikram, et al.
Published: (2026)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
by: Zhu, Ziyi, et al.
Published: (2026)
by: Zhu, Ziyi, et al.
Published: (2026)
Approximating Simplet Frequency Distribution for Simplicial Complexes
by: Beigy, Hamid, et al.
Published: (2024)
by: Beigy, Hamid, et al.
Published: (2024)
RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
by: Wei, Tianjun, et al.
Published: (2025)
by: Wei, Tianjun, et al.
Published: (2025)
Quantitative LLM Judges
by: Sahoo, Aishwarya, et al.
Published: (2025)
by: Sahoo, Aishwarya, et al.
Published: (2025)
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
by: Sutawika, Lintang, et al.
Published: (2026)
by: Sutawika, Lintang, et al.
Published: (2026)
Mina: A Multilingual LLM-Powered Legal Assistant Agent for Bangladesh for Empowering Access to Justice
by: Wasi, Azmine Toushik, et al.
Published: (2025)
by: Wasi, Azmine Toushik, et al.
Published: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
Can LLM be a Personalized Judge?
by: Dong, Yijiang River, et al.
Published: (2024)
by: Dong, Yijiang River, et al.
Published: (2024)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
by: Lee, Dongryeol, et al.
Published: (2026)
by: Lee, Dongryeol, et al.
Published: (2026)
ALOHA: Empowering Multilingual Agent for University Orientation with Hierarchical Retrieval
by: Tao, Mingxu, et al.
Published: (2025)
by: Tao, Mingxu, et al.
Published: (2025)
Evaluating Scoring Bias in LLM-as-a-Judge
by: Li, Qingquan, et al.
Published: (2025)
by: Li, Qingquan, et al.
Published: (2025)
The Necessity of Setting Temperature in LLM-as-a-Judge
by: Li, Lujun, et al.
Published: (2026)
by: Li, Lujun, et al.
Published: (2026)
Self-Preference Bias in LLM-as-a-Judge
by: Wataoka, Koki, et al.
Published: (2024)
by: Wataoka, Koki, et al.
Published: (2024)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
by: Huang, Hui, et al.
Published: (2024)
by: Huang, Hui, et al.
Published: (2024)
Smart Audit System Empowered by LLM
by: Yao, Xu, et al.
Published: (2024)
by: Yao, Xu, et al.
Published: (2024)
A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
by: Guo, Jiale, et al.
Published: (2025)
by: Guo, Jiale, et al.
Published: (2025)
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
by: Schroeder, Kayla, et al.
Published: (2024)
by: Schroeder, Kayla, et al.
Published: (2024)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
by: Wang, Yidong, et al.
Published: (2025)
by: Wang, Yidong, et al.
Published: (2025)
Similar Items
-
Zero-Shot Learning and Key Points Are All You Need for Automated Fact-Checking
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2024) -
Gap-Filling Prompting Enhances Code-Assisted Mathematical Reasoning
by: Mohammadkhani, Mohammad Ghiasvand
Published: (2024) -
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
by: Mohammadkhani, Ali Ghiasvand
Published: (2024) -
E2TP: Element to Tuple Prompting Improves Aspect Sentiment Tuple Prediction
by: Mohammadkhani, Mohammad Ghiasvand, et al.
Published: (2024) -
Consistency Training by Synthetic Question Generation for Conversational Question Answering
by: Hemati, Hamed Hematian, et al.
Published: (2024)