Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tang, Yuqi, Feng, Kehua, Wang, Yunfeng, Chen, Zhiwen, Lv, Chengfei, Yu, Gang, Zhang, Qiang, Ding, Keyan, Chen, Huajun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914234669989888
author Tang, Yuqi
Feng, Kehua
Wang, Yunfeng
Chen, Zhiwen
Lv, Chengfei
Yu, Gang
Zhang, Qiang
Ding, Keyan
Chen, Huajun
author_facet Tang, Yuqi
Feng, Kehua
Wang, Yunfeng
Chen, Zhiwen
Lv, Chengfei
Yu, Gang
Zhang, Qiang
Ding, Keyan
Chen, Huajun
contents Evaluating the conversational abilities of large language models (LLMs) remains a challenging task. Current mainstream approaches primarily rely on the "LLM-as-a-judge" paradigm, where an LLM is prompted to serve as an evaluator to assess dialogue quality. However, such methods often suffer from various biases, which undermine the reliability and consistency of the evaluation results. To mitigate these biases, recent methods employ multiple LLMs as judges and aggregate their judgments to select the optimal assessment. Although effective, this multi-judge approach incurs significant computational overhead during inference. In this paper, we propose an efficient dialogue evaluator that captures the collective wisdom of multiple LLM judges by aggregating their preference knowledge into a single model. Our approach preserves the advantages of diverse multi-judge feedback while drastically reducing the evaluation cost, enabling fast, flexible, and fine-grained dialogue quality assessment. Extensive experiments on seven single rating and pairwise comparison dialogue evaluation benchmarks demonstrate that our method outperforms existing baselines across diverse scenarios, showcasing its efficiency and robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00454
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges
Tang, Yuqi
Feng, Kehua
Wang, Yunfeng
Chen, Zhiwen
Lv, Chengfei
Yu, Gang
Zhang, Qiang
Ding, Keyan
Chen, Huajun
Computation and Language
Evaluating the conversational abilities of large language models (LLMs) remains a challenging task. Current mainstream approaches primarily rely on the "LLM-as-a-judge" paradigm, where an LLM is prompted to serve as an evaluator to assess dialogue quality. However, such methods often suffer from various biases, which undermine the reliability and consistency of the evaluation results. To mitigate these biases, recent methods employ multiple LLMs as judges and aggregate their judgments to select the optimal assessment. Although effective, this multi-judge approach incurs significant computational overhead during inference. In this paper, we propose an efficient dialogue evaluator that captures the collective wisdom of multiple LLM judges by aggregating their preference knowledge into a single model. Our approach preserves the advantages of diverse multi-judge feedback while drastically reducing the evaluation cost, enabling fast, flexible, and fine-grained dialogue quality assessment. Extensive experiments on seven single rating and pairwise comparison dialogue evaluation benchmarks demonstrate that our method outperforms existing baselines across diverse scenarios, showcasing its efficiency and robustness.
title Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges
topic Computation and Language
url https://arxiv.org/abs/2508.00454