Saved in:
Bibliographic Details
Main Authors: Jeong, Hawon, Park, ChaeHun, Hong, Jimin, Lee, Hojoon, Choo, Jaegul
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2406.12319
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915249133715456
author Jeong, Hawon
Park, ChaeHun
Hong, Jimin
Lee, Hojoon
Choo, Jaegul
author_facet Jeong, Hawon
Park, ChaeHun
Hong, Jimin
Lee, Hojoon
Choo, Jaegul
contents As large language models (LLMs) are increasingly used as evaluators for natural language generation tasks, ensuring unbiased assessments is essential. However, LLM evaluators often display biased preferences, such as favoring verbosity and authoritative tones. Our empirical analysis reveals that these biases are exacerbated in pairwise evaluation, where LLMs directly compare two outputs and easily prioritize superficial attributes. In contrast, pointwise evaluation, which assesses outputs independently, is less susceptible to such bias because each output is judged in isolation. To address the limitations of the pairwise evaluation, we introduce a novel evaluation method, PRePair, which integrates pointwise reasoning within a pairwise framework. PRePair effectively alleviates biased preference, improving performance on the adversarial benchmark (LLMBar) while outperforming pointwise evaluation on the standard benchmark (MT-Bench).
format Preprint
id arxiv_https___arxiv_org_abs_2406_12319
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Comparative Trap: Pairwise Comparisons Amplifies Biased Preferences of LLM Evaluators
Jeong, Hawon
Park, ChaeHun
Hong, Jimin
Lee, Hojoon
Choo, Jaegul
Computation and Language
As large language models (LLMs) are increasingly used as evaluators for natural language generation tasks, ensuring unbiased assessments is essential. However, LLM evaluators often display biased preferences, such as favoring verbosity and authoritative tones. Our empirical analysis reveals that these biases are exacerbated in pairwise evaluation, where LLMs directly compare two outputs and easily prioritize superficial attributes. In contrast, pointwise evaluation, which assesses outputs independently, is less susceptible to such bias because each output is judged in isolation. To address the limitations of the pairwise evaluation, we introduce a novel evaluation method, PRePair, which integrates pointwise reasoning within a pairwise framework. PRePair effectively alleviates biased preference, improving performance on the adversarial benchmark (LLMBar) while outperforming pointwise evaluation on the standard benchmark (MT-Bench).
title The Comparative Trap: Pairwise Comparisons Amplifies Biased Preferences of LLM Evaluators
topic Computation and Language
url https://arxiv.org/abs/2406.12319