Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khalifa, Muhammad, Logeswaran, Lajanugen, Kim, Jaekyeom, Sohn, Sungryull, Zhang, Yunxiang, Lee, Moontae, Peng, Hao, Wang, Lu, Lee, Honglak
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917216586301440
author Khalifa, Muhammad
Logeswaran, Lajanugen
Kim, Jaekyeom
Sohn, Sungryull
Zhang, Yunxiang
Lee, Moontae
Peng, Hao
Wang, Lu
Lee, Honglak
author_facet Khalifa, Muhammad
Logeswaran, Lajanugen
Kim, Jaekyeom
Sohn, Sungryull
Zhang, Yunxiang
Lee, Moontae
Peng, Hao
Wang, Lu
Lee, Honglak
contents Large language models (LLMs) are increasingly used as judges to evaluate agent performance, particularly in non-verifiable settings where judgments rely on agent trajectories including chain-of-thought (CoT) reasoning. This paradigm implicitly assumes that the agent's CoT faithfully reflects both its internal reasoning and the underlying environment state. We show this assumption is brittle: LLM judges are highly susceptible to manipulation of agent reasoning traces. By systematically rewriting agent CoTs while holding actions and observations fixed, we demonstrate that manipulated reasoning alone can inflate false positive rates of state-of-the-art VLM judges by up to 90% across 800 trajectories spanning diverse web tasks. We study manipulation strategies spanning style-based approaches that alter only the presentation of reasoning and content-based approaches that fabricate signals of task progress, and find that content-based manipulations are consistently more effective. We evaluate prompting-based techniques and scaling judge-time compute, which reduce but do not fully eliminate susceptibility to manipulation. Our findings reveal a fundamental vulnerability in LLM-based evaluation and highlight the need for judging mechanisms that verify reasoning claims against observable evidence.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14691
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
Khalifa, Muhammad
Logeswaran, Lajanugen
Kim, Jaekyeom
Sohn, Sungryull
Zhang, Yunxiang
Lee, Moontae
Peng, Hao
Wang, Lu
Lee, Honglak
Artificial Intelligence
Computation and Language
Large language models (LLMs) are increasingly used as judges to evaluate agent performance, particularly in non-verifiable settings where judgments rely on agent trajectories including chain-of-thought (CoT) reasoning. This paradigm implicitly assumes that the agent's CoT faithfully reflects both its internal reasoning and the underlying environment state. We show this assumption is brittle: LLM judges are highly susceptible to manipulation of agent reasoning traces. By systematically rewriting agent CoTs while holding actions and observations fixed, we demonstrate that manipulated reasoning alone can inflate false positive rates of state-of-the-art VLM judges by up to 90% across 800 trajectories spanning diverse web tasks. We study manipulation strategies spanning style-based approaches that alter only the presentation of reasoning and content-based approaches that fabricate signals of task progress, and find that content-based manipulations are consistently more effective. We evaluate prompting-based techniques and scaling judge-time compute, which reduce but do not fully eliminate susceptibility to manipulation. Our findings reveal a fundamental vulnerability in LLM-based evaluation and highlight the need for judging mechanisms that verify reasoning claims against observable evidence.
title Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.14691