Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hu, Zhengyu, Zhang, Jieyu, Xiong, Zhihan, Ratner, Alexander, Ding, Kaize, Krishna, Ranjay
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917179111243776
author Hu, Zhengyu
Zhang, Jieyu
Xiong, Zhihan
Ratner, Alexander
Ding, Kaize
Krishna, Ranjay
author_facet Hu, Zhengyu
Zhang, Jieyu
Xiong, Zhihan
Ratner, Alexander
Ding, Kaize
Krishna, Ranjay
contents Despite the remarkable success of Large Language Models (LLMs), evaluating their outputs' quality regarding preference remains a critical challenge. While existing works usually leverage a strong LLM as the judge for comparing LLMs' response pairwisely, such a single-evaluator approach is vulnerable to cyclic preference, i.e., output A is better than B, B than C, but C is better than A, causing contradictory evaluation results. To address this, we introduce PGED (Preference Graph Ensemble and Denoising), a novel approach that leverages multiple model-based evaluators to construct preference graphs, and then ensembles and denoises these graphs for acyclic, non-contradictory evaluation results. We provide theoretical guarantees for our framework, demonstrating its efficacy in recovering the ground truth preference structure. Extensive experiments on ten benchmarks demonstrate PGED's superiority in three applications: 1) model ranking for evaluation, 2) response selection for test-time scaling, and 3) data selection for model fine-tuning. Notably, PGED combines small LLM evaluators (e.g., Llama3-8B, Mistral-7B, Qwen2-7B) to outperform strong ones (e.g., Qwen2-72B), showcasing its effectiveness in enhancing evaluation reliability and improving model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12869
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
Hu, Zhengyu
Zhang, Jieyu
Xiong, Zhihan
Ratner, Alexander
Ding, Kaize
Krishna, Ranjay
Computation and Language
Artificial Intelligence
Machine Learning
Despite the remarkable success of Large Language Models (LLMs), evaluating their outputs' quality regarding preference remains a critical challenge. While existing works usually leverage a strong LLM as the judge for comparing LLMs' response pairwisely, such a single-evaluator approach is vulnerable to cyclic preference, i.e., output A is better than B, B than C, but C is better than A, causing contradictory evaluation results. To address this, we introduce PGED (Preference Graph Ensemble and Denoising), a novel approach that leverages multiple model-based evaluators to construct preference graphs, and then ensembles and denoises these graphs for acyclic, non-contradictory evaluation results. We provide theoretical guarantees for our framework, demonstrating its efficacy in recovering the ground truth preference structure. Extensive experiments on ten benchmarks demonstrate PGED's superiority in three applications: 1) model ranking for evaluation, 2) response selection for test-time scaling, and 3) data selection for model fine-tuning. Notably, PGED combines small LLM evaluators (e.g., Llama3-8B, Mistral-7B, Qwen2-7B) to outperform strong ones (e.g., Qwen2-72B), showcasing its effectiveness in enhancing evaluation reliability and improving model performance.
title Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.12869