Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yangsibo, Nasr, Milad, Angelopoulos, Anastasios, Carlini, Nicholas, Chiang, Wei-Lin, Choquette-Choo, Christopher A., Ippolito, Daphne, Jagielski, Matthew, Lee, Katherine, Liu, Ken Ziyu, Stoica, Ion, Tramer, Florian, Zhang, Chiyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915101178593280
author Huang, Yangsibo
Nasr, Milad
Angelopoulos, Anastasios
Carlini, Nicholas
Chiang, Wei-Lin
Choquette-Choo, Christopher A.
Ippolito, Daphne
Jagielski, Matthew
Lee, Katherine
Liu, Ken Ziyu
Stoica, Ion
Tramer, Florian
Zhang, Chiyuan
author_facet Huang, Yangsibo
Nasr, Milad
Angelopoulos, Anastasios
Carlini, Nicholas
Chiang, Wei-Lin
Choquette-Choo, Christopher A.
Ippolito, Daphne
Jagielski, Matthew
Lee, Katherine
Liu, Ken Ziyu
Stoica, Ion
Tramer, Florian
Zhang, Chiyuan
contents It is now common to evaluate Large Language Models (LLMs) by having humans manually vote to evaluate model outputs, in contrast to typical benchmarks that evaluate knowledge or skill at some particular task. Chatbot Arena, the most popular benchmark of this type, ranks models by asking users to select the better response between two randomly selected models (without revealing which model was responsible for the generations). These platforms are widely trusted as a fair and accurate measure of LLM capabilities. In this paper, we show that if bot protection and other defenses are not implemented, these voting-based benchmarks are potentially vulnerable to adversarial manipulation. Specifically, we show that an attacker can alter the leaderboard (to promote their favorite model or demote competitors) at the cost of roughly a thousand votes (verified in a simulated, offline version of Chatbot Arena). Our attack consists of two steps: first, we show how an attacker can determine which model was used to generate a given reply with more than $95\%$ accuracy; and then, the attacker can use this information to consistently vote for (or against) a target model. Working with the Chatbot Arena developers, we identify, propose, and implement mitigations to improve the robustness of Chatbot Arena against adversarial manipulation, which, based on our analysis, substantially increases the cost of such attacks. Some of these defenses were present before our collaboration, such as bot protection with Cloudflare, malicious user detection, and rate limiting. Others, including reCAPTCHA and login are being integrated to strengthen the security in Chatbot Arena.
format Preprint
id arxiv_https___arxiv_org_abs_2501_07493
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
Huang, Yangsibo
Nasr, Milad
Angelopoulos, Anastasios
Carlini, Nicholas
Chiang, Wei-Lin
Choquette-Choo, Christopher A.
Ippolito, Daphne
Jagielski, Matthew
Lee, Katherine
Liu, Ken Ziyu
Stoica, Ion
Tramer, Florian
Zhang, Chiyuan
Machine Learning
Cryptography and Security
It is now common to evaluate Large Language Models (LLMs) by having humans manually vote to evaluate model outputs, in contrast to typical benchmarks that evaluate knowledge or skill at some particular task. Chatbot Arena, the most popular benchmark of this type, ranks models by asking users to select the better response between two randomly selected models (without revealing which model was responsible for the generations). These platforms are widely trusted as a fair and accurate measure of LLM capabilities. In this paper, we show that if bot protection and other defenses are not implemented, these voting-based benchmarks are potentially vulnerable to adversarial manipulation. Specifically, we show that an attacker can alter the leaderboard (to promote their favorite model or demote competitors) at the cost of roughly a thousand votes (verified in a simulated, offline version of Chatbot Arena). Our attack consists of two steps: first, we show how an attacker can determine which model was used to generate a given reply with more than $95\%$ accuracy; and then, the attacker can use this information to consistently vote for (or against) a target model. Working with the Chatbot Arena developers, we identify, propose, and implement mitigations to improve the robustness of Chatbot Arena against adversarial manipulation, which, based on our analysis, substantially increases the cost of such attacks. Some of these defenses were present before our collaboration, such as bot protection with Cloudflare, malicious user detection, and rate limiting. Others, including reCAPTCHA and login are being integrated to strengthen the security in Chatbot Arena.
title Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2501.07493