Saved in:
Bibliographic Details
Main Authors: Hu, Zhengyu, Song, Linxin, Zhang, Jieyu, Xiao, Zheyuan, Wang, Tianfu, Chen, Zhengyu, Yuan, Nicholas Jing, Lian, Jianxun, Ding, Kaize, Xiong, Hui
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.01085
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914020863246336
author Hu, Zhengyu
Song, Linxin
Zhang, Jieyu
Xiao, Zheyuan
Wang, Tianfu
Chen, Zhengyu
Yuan, Nicholas Jing
Lian, Jianxun
Ding, Kaize
Xiong, Hui
author_facet Hu, Zhengyu
Song, Linxin
Zhang, Jieyu
Xiao, Zheyuan
Wang, Tianfu
Chen, Zhengyu
Yuan, Nicholas Jing
Lian, Jianxun
Ding, Kaize
Xiong, Hui
contents The use of large language models (LLMs) as judges, particularly in preference comparisons, has become widespread, but this reveals a notable bias towards longer responses, undermining the reliability of such evaluations. To better understand such bias, we propose to decompose the preference evaluation metric, specifically the win rate, into two key components: desirability and information mass, where the former is length-independent and related to trustworthiness such as correctness, toxicity, and consistency, and the latter is length-dependent and represents the amount of information in the response. We empirically demonstrated the decomposition through controlled experiments and found that response length impacts evaluations by influencing information mass. To derive a reliable evaluation metric that assesses content quality without being confounded by response length, we propose AdapAlpaca, a simple yet effective adjustment to win rate measurement. Specifically, AdapAlpaca ensures a fair comparison of response quality by aligning the lengths of reference and test model responses under equivalent length intervals.
format Preprint
id arxiv_https___arxiv_org_abs_2407_01085
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Explaining Length Bias in LLM-Based Preference Evaluations
Hu, Zhengyu
Song, Linxin
Zhang, Jieyu
Xiao, Zheyuan
Wang, Tianfu
Chen, Zhengyu
Yuan, Nicholas Jing
Lian, Jianxun
Ding, Kaize
Xiong, Hui
Machine Learning
Computation and Language
The use of large language models (LLMs) as judges, particularly in preference comparisons, has become widespread, but this reveals a notable bias towards longer responses, undermining the reliability of such evaluations. To better understand such bias, we propose to decompose the preference evaluation metric, specifically the win rate, into two key components: desirability and information mass, where the former is length-independent and related to trustworthiness such as correctness, toxicity, and consistency, and the latter is length-dependent and represents the amount of information in the response. We empirically demonstrated the decomposition through controlled experiments and found that response length impacts evaluations by influencing information mass. To derive a reliable evaluation metric that assesses content quality without being confounded by response length, we propose AdapAlpaca, a simple yet effective adjustment to win rate measurement. Specifically, AdapAlpaca ensures a fair comparison of response quality by aligning the lengths of reference and test model responses under equivalent length intervals.
title Explaining Length Bias in LLM-Based Preference Evaluations
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2407.01085