Can You Trick the Grader? Adversarial Persuasion of LLM Judges

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hwang, Yerin, Lee, Dongryeol, Kang, Taegwan, Kim, Yongil, Jung, Kyomin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915439603351552
author Hwang, Yerin
Lee, Dongryeol
Kang, Taegwan
Kim, Yongil
Jung, Kyomin
author_facet Hwang, Yerin
Lee, Dongryeol
Kang, Taegwan
Kim, Yongil
Jung, Kyomin
contents As large language models take on growing roles as automated evaluators in practical settings, a critical question arises: Can individuals persuade an LLM judge to assign unfairly high scores? This study is the first to reveal that strategically embedded persuasive language can bias LLM judges when scoring mathematical reasoning tasks, where correctness should be independent of stylistic variation. Grounded in Aristotle's rhetorical principles, we formalize seven persuasion techniques (Majority, Consistency, Flattery, Reciprocity, Pity, Authority, Identity) and embed them into otherwise identical responses. Across six math benchmarks, we find that persuasive language leads LLM judges to assign inflated scores to incorrect solutions, by up to 8% on average, with Consistency causing the most severe distortion. Notably, increasing model size does not substantially mitigate this vulnerability. Further analysis demonstrates that combining multiple persuasion techniques amplifies the bias, and pairwise evaluation is likewise susceptible. Moreover, the persuasive effect persists under counter prompting strategies, highlighting a critical vulnerability in LLM-as-a-Judge pipelines and underscoring the need for robust defenses against persuasion-based attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can You Trick the Grader? Adversarial Persuasion of LLM Judges
Hwang, Yerin
Lee, Dongryeol
Kang, Taegwan
Kim, Yongil
Jung, Kyomin
Computation and Language
As large language models take on growing roles as automated evaluators in practical settings, a critical question arises: Can individuals persuade an LLM judge to assign unfairly high scores? This study is the first to reveal that strategically embedded persuasive language can bias LLM judges when scoring mathematical reasoning tasks, where correctness should be independent of stylistic variation. Grounded in Aristotle's rhetorical principles, we formalize seven persuasion techniques (Majority, Consistency, Flattery, Reciprocity, Pity, Authority, Identity) and embed them into otherwise identical responses. Across six math benchmarks, we find that persuasive language leads LLM judges to assign inflated scores to incorrect solutions, by up to 8% on average, with Consistency causing the most severe distortion. Notably, increasing model size does not substantially mitigate this vulnerability. Further analysis demonstrates that combining multiple persuasion techniques amplifies the bias, and pairwise evaluation is likewise susceptible. Moreover, the persuasive effect persists under counter prompting strategies, highlighting a critical vulnerability in LLM-as-a-Judge pipelines and underscoring the need for robust defenses against persuasion-based attacks.
title Can You Trick the Grader? Adversarial Persuasion of LLM Judges
topic Computation and Language
url https://arxiv.org/abs/2508.07805