Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Khalifa, Muhammad, Logeswaran, Lajanugen, Kim, Jaekyeom, Sohn, Sungryull, Zhang, Yunxiang, Lee, Moontae, Peng, Hao, Wang, Lu, Lee, Honglak |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
by: Khalifa, Muhammad, et al.
Published: (2023)
by: Khalifa, Muhammad, et al.
Published: (2023)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
by: Kim, Jaekyeom, et al.
Published: (2024)
by: Kim, Jaekyeom, et al.
Published: (2024)
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
by: Zhang, Yunxiang, et al.
Published: (2024)
by: Zhang, Yunxiang, et al.
Published: (2024)
Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
by: Logeswaran, Lajanugen, et al.
Published: (2026)
by: Logeswaran, Lajanugen, et al.
Published: (2026)
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025)
by: Khalifa, Muhammad, et al.
Published: (2025)
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
by: Fu, Yao, et al.
Published: (2024)
by: Fu, Yao, et al.
Published: (2024)
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
by: Zhang, Yunxiang, et al.
Published: (2025)
by: Zhang, Yunxiang, et al.
Published: (2025)
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
by: Jang, Yunseok, et al.
Published: (2025)
by: Jang, Yunseok, et al.
Published: (2025)
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
by: Shen, Siqi, et al.
Published: (2025)
by: Shen, Siqi, et al.
Published: (2025)
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
by: Shen, Siqi, et al.
Published: (2024)
by: Shen, Siqi, et al.
Published: (2024)
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Interactive and Expressive Code-Augmented Planning with Large Language Models
by: Liu, Anthony Z., et al.
Published: (2024)
by: Liu, Anthony Z., et al.
Published: (2024)
SPRIG: Improving Large Language Model Performance by System Prompt Optimization
by: Zhang, Lechen, et al.
Published: (2024)
by: Zhang, Lechen, et al.
Published: (2024)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
by: Zheng, Mingqian, et al.
Published: (2023)
by: Zheng, Mingqian, et al.
Published: (2023)
Visual Test-time Scaling for GUI Agent Grounding
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Cross-Lingual Prompt Steerability: Towards Accurate and Robust LLM Behavior across Languages
by: Zhang, Lechen, et al.
Published: (2025)
by: Zhang, Lechen, et al.
Published: (2025)
Chain-of-Thought Unfaithfulness as Disguised Accuracy
by: Bentham, Oliver, et al.
Published: (2024)
by: Bentham, Oliver, et al.
Published: (2024)
You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments
by: Shu, Bangzhao, et al.
Published: (2023)
by: Shu, Bangzhao, et al.
Published: (2023)
Do Not Trust Licenses You See: Dataset Compliance Requires Massive-Scale AI-Powered Lifecycle Tracing
by: Kim, Jaekyeom, et al.
Published: (2025)
by: Kim, Jaekyeom, et al.
Published: (2025)
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
by: Han, Janghoon, et al.
Published: (2025)
by: Han, Janghoon, et al.
Published: (2025)
Multi-round, Chain-of-thought Post-editing for Unfaithful Summaries
by: Lee, Yi-Hui, et al.
Published: (2025)
by: Lee, Yi-Hui, et al.
Published: (2025)
Source-Aware Training Enables Knowledge Attribution in Language Models
by: Khalifa, Muhammad, et al.
Published: (2024)
by: Khalifa, Muhammad, et al.
Published: (2024)
Can Separators Improve Chain-of-Thought Prompting?
by: Park, Yoonjeong, et al.
Published: (2024)
by: Park, Yoonjeong, et al.
Published: (2024)
Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?
by: Lewis-Lim, Samuel, et al.
Published: (2025)
by: Lewis-Lim, Samuel, et al.
Published: (2025)
If You Can't Use Them, Recycle Them: Optimizing Merging at Scale Mitigates Performance Tradeoffs
by: Khalifa, Muhammad, et al.
Published: (2024)
by: Khalifa, Muhammad, et al.
Published: (2024)
How to Correctly Report LLM-as-a-Judge Evaluations
by: Lee, Chungpa, et al.
Published: (2025)
by: Lee, Chungpa, et al.
Published: (2025)
Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation
by: Koupaee, Mahnaz, et al.
Published: (2025)
by: Koupaee, Mahnaz, et al.
Published: (2025)
TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge
by: Chiang, Cheng-Han, et al.
Published: (2025)
by: Chiang, Cheng-Han, et al.
Published: (2025)
Pre-trained Language Models Return Distinguishable Probability Distributions to Unfaithfully Hallucinated Texts
by: Cha, Taehun, et al.
Published: (2024)
by: Cha, Taehun, et al.
Published: (2024)
When Is Enough Not Enough? Illusory Completion in Search Agents
by: Ko, Dayoon, et al.
Published: (2026)
by: Ko, Dayoon, et al.
Published: (2026)
On Many-Shot In-Context Learning for Long-Context Evaluation
by: Zou, Kaijian, et al.
Published: (2024)
by: Zou, Kaijian, et al.
Published: (2024)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
by: Hwang, Yerin, et al.
Published: (2025)
by: Hwang, Yerin, et al.
Published: (2025)
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
by: Zhang, Yunxiang, et al.
Published: (2025)
by: Zhang, Yunxiang, et al.
Published: (2025)
Dissociation of Faithful and Unfaithful Reasoning in LLMs
by: Yee, Evelyn, et al.
Published: (2024)
by: Yee, Evelyn, et al.
Published: (2024)
Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts
by: Chen, Yi-Chang, et al.
Published: (2026)
by: Chen, Yi-Chang, et al.
Published: (2026)
Faithful Logical Reasoning via Symbolic Chain-of-Thought
by: Xu, Jundong, et al.
Published: (2024)
by: Xu, Jundong, et al.
Published: (2024)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
Shifting from Ranking to Set Selection for Retrieval Augmented Generation
by: Lee, Dahyun, et al.
Published: (2025)
by: Lee, Dahyun, et al.
Published: (2025)
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
by: Zhang, Yunxiang, et al.
Published: (2025)
by: Zhang, Yunxiang, et al.
Published: (2025)
FactCheckmate: Preemptively Detecting and Mitigating Hallucinations in LMs
by: Alnuhait, Deema, et al.
Published: (2024)
by: Alnuhait, Deema, et al.
Published: (2024)
Similar Items
-
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
by: Khalifa, Muhammad, et al.
Published: (2023) -
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
by: Kim, Jaekyeom, et al.
Published: (2024) -
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
by: Zhang, Yunxiang, et al.
Published: (2024) -
Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
by: Logeswaran, Lajanugen, et al.
Published: (2026) -
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025)