Challenges in Trustworthy Human Evaluation of Chatbots

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Wenting, Rush, Alexander M., Goyal, Tanya
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909417588391936
author Zhao, Wenting
Rush, Alexander M.
Goyal, Tanya
author_facet Zhao, Wenting
Rush, Alexander M.
Goyal, Tanya
contents Open community-driven platforms like Chatbot Arena that collect user preference data from site visitors have gained a reputation as one of the most trustworthy publicly available benchmarks for LLM performance. While now standard, it is tricky to implement effective guardrails to collect high-quality annotations from humans. In this paper, we demonstrate that three sources of bad annotations, both malicious and otherwise, can corrupt the reliability of open leaderboard rankings. In particular, we show that only 10\% of poor quality votes by apathetic (site visitors not appropriately incentivized to give correct votes) or adversarial (bad actors seeking to inflate the ranking of a target model) annotators can change the rankings of models by up to 5 places on the leaderboard. Finally, we discuss open challenges in ensuring high-quality human annotations.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04363
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Challenges in Trustworthy Human Evaluation of Chatbots
Zhao, Wenting
Rush, Alexander M.
Goyal, Tanya
Human-Computer Interaction
Open community-driven platforms like Chatbot Arena that collect user preference data from site visitors have gained a reputation as one of the most trustworthy publicly available benchmarks for LLM performance. While now standard, it is tricky to implement effective guardrails to collect high-quality annotations from humans. In this paper, we demonstrate that three sources of bad annotations, both malicious and otherwise, can corrupt the reliability of open leaderboard rankings. In particular, we show that only 10\% of poor quality votes by apathetic (site visitors not appropriately incentivized to give correct votes) or adversarial (bad actors seeking to inflate the ranking of a target model) annotators can change the rankings of models by up to 5 places on the leaderboard. Finally, we discuss open challenges in ensuring high-quality human annotations.
title Challenges in Trustworthy Human Evaluation of Chatbots
topic Human-Computer Interaction
url https://arxiv.org/abs/2412.04363