An Empirical Analysis on Large Language Models in Debate Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xinyi, Liu, Pinxin, He, Hangfeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929372832727040
author Liu, Xinyi
Liu, Pinxin
He, Hangfeng
author_facet Liu, Xinyi
Liu, Pinxin
He, Hangfeng
contents In this study, we investigate the capabilities and inherent biases of advanced large language models (LLMs) such as GPT-3.5 and GPT-4 in the context of debate evaluation. We discover that LLM's performance exceeds humans and surpasses the performance of state-of-the-art methods fine-tuned on extensive datasets in debate evaluation. We additionally explore and analyze biases present in LLMs, including positional bias, lexical bias, order bias, which may affect their evaluative judgments. Our findings reveal a consistent bias in both GPT-3.5 and GPT-4 towards the second candidate response presented, attributed to prompt design. We also uncover lexical biases in both GPT-3.5 and GPT-4, especially when label sets carry connotations such as numerical or sequential, highlighting the critical need for careful label verbalizer selection in prompt design. Additionally, our analysis indicates a tendency of both models to favor the debate's concluding side as the winner, suggesting an end-of-discussion bias.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00050
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Empirical Analysis on Large Language Models in Debate Evaluation
Liu, Xinyi
Liu, Pinxin
He, Hangfeng
Computation and Language
Artificial Intelligence
In this study, we investigate the capabilities and inherent biases of advanced large language models (LLMs) such as GPT-3.5 and GPT-4 in the context of debate evaluation. We discover that LLM's performance exceeds humans and surpasses the performance of state-of-the-art methods fine-tuned on extensive datasets in debate evaluation. We additionally explore and analyze biases present in LLMs, including positional bias, lexical bias, order bias, which may affect their evaluative judgments. Our findings reveal a consistent bias in both GPT-3.5 and GPT-4 towards the second candidate response presented, attributed to prompt design. We also uncover lexical biases in both GPT-3.5 and GPT-4, especially when label sets carry connotations such as numerical or sequential, highlighting the critical need for careful label verbalizer selection in prompt design. Additionally, our analysis indicates a tendency of both models to favor the debate's concluding side as the winner, suggesting an end-of-discussion bias.
title An Empirical Analysis on Large Language Models in Debate Evaluation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.00050