Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.01085 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914020863246336 |
|---|---|
| author | Hu, Zhengyu Song, Linxin Zhang, Jieyu Xiao, Zheyuan Wang, Tianfu Chen, Zhengyu Yuan, Nicholas Jing Lian, Jianxun Ding, Kaize Xiong, Hui |
| author_facet | Hu, Zhengyu Song, Linxin Zhang, Jieyu Xiao, Zheyuan Wang, Tianfu Chen, Zhengyu Yuan, Nicholas Jing Lian, Jianxun Ding, Kaize Xiong, Hui |
| contents | The use of large language models (LLMs) as judges, particularly in preference comparisons, has become widespread, but this reveals a notable bias towards longer responses, undermining the reliability of such evaluations. To better understand such bias, we propose to decompose the preference evaluation metric, specifically the win rate, into two key components: desirability and information mass, where the former is length-independent and related to trustworthiness such as correctness, toxicity, and consistency, and the latter is length-dependent and represents the amount of information in the response. We empirically demonstrated the decomposition through controlled experiments and found that response length impacts evaluations by influencing information mass. To derive a reliable evaluation metric that assesses content quality without being confounded by response length, we propose AdapAlpaca, a simple yet effective adjustment to win rate measurement. Specifically, AdapAlpaca ensures a fair comparison of response quality by aligning the lengths of reference and test model responses under equivalent length intervals. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_01085 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Explaining Length Bias in LLM-Based Preference Evaluations Hu, Zhengyu Song, Linxin Zhang, Jieyu Xiao, Zheyuan Wang, Tianfu Chen, Zhengyu Yuan, Nicholas Jing Lian, Jianxun Ding, Kaize Xiong, Hui Machine Learning Computation and Language The use of large language models (LLMs) as judges, particularly in preference comparisons, has become widespread, but this reveals a notable bias towards longer responses, undermining the reliability of such evaluations. To better understand such bias, we propose to decompose the preference evaluation metric, specifically the win rate, into two key components: desirability and information mass, where the former is length-independent and related to trustworthiness such as correctness, toxicity, and consistency, and the latter is length-dependent and represents the amount of information in the response. We empirically demonstrated the decomposition through controlled experiments and found that response length impacts evaluations by influencing information mass. To derive a reliable evaluation metric that assesses content quality without being confounded by response length, we propose AdapAlpaca, a simple yet effective adjustment to win rate measurement. Specifically, AdapAlpaca ensures a fair comparison of response quality by aligning the lengths of reference and test model responses under equivalent length intervals. |
| title | Explaining Length Bias in LLM-Based Preference Evaluations |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2407.01085 |