Checklist Engineering Empowers Multilingual LLM Judges

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mohammadkhani, Mohammad Ghiasvand, Beigy, Hamid
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911079103201280
author Mohammadkhani, Mohammad Ghiasvand
Beigy, Hamid
author_facet Mohammadkhani, Mohammad Ghiasvand
Beigy, Hamid
contents Automated text evaluation has long been a central issue in Natural Language Processing (NLP). Recently, the field has shifted toward using Large Language Models (LLMs) as evaluators-a trend known as the LLM-as-a-Judge paradigm. While promising and easily adaptable across tasks, this approach has seen limited exploration in multilingual contexts. Existing multilingual studies often rely on proprietary models or require extensive training data for fine-tuning, raising concerns about cost, time, and efficiency. In this paper, we propose Checklist Engineering based LLM-as-a-Judge (CE-Judge), a training-free framework that uses checklist intuition for multilingual evaluation with an open-source model. Experiments across multiple languages and three benchmark datasets, under both pointwise and pairwise settings, show that our method generally surpasses the baselines and performs on par with the GPT-4o model.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06774
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Checklist Engineering Empowers Multilingual LLM Judges
Mohammadkhani, Mohammad Ghiasvand
Beigy, Hamid
Computation and Language
Automated text evaluation has long been a central issue in Natural Language Processing (NLP). Recently, the field has shifted toward using Large Language Models (LLMs) as evaluators-a trend known as the LLM-as-a-Judge paradigm. While promising and easily adaptable across tasks, this approach has seen limited exploration in multilingual contexts. Existing multilingual studies often rely on proprietary models or require extensive training data for fine-tuning, raising concerns about cost, time, and efficiency. In this paper, we propose Checklist Engineering based LLM-as-a-Judge (CE-Judge), a training-free framework that uses checklist intuition for multilingual evaluation with an open-source model. Experiments across multiple languages and three benchmark datasets, under both pointwise and pairwise settings, show that our method generally surpasses the baselines and performs on par with the GPT-4o model.
title Checklist Engineering Empowers Multilingual LLM Judges
topic Computation and Language
url https://arxiv.org/abs/2507.06774