Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Dongryeol, Hwang, Yerin, Kang, Taegwan, Lee, Minwoo, Chae, Younhyung, Jung, Kyomin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912818257723392
author Lee, Dongryeol
Hwang, Yerin
Kang, Taegwan
Lee, Minwoo
Chae, Younhyung
Jung, Kyomin
author_facet Lee, Dongryeol
Hwang, Yerin
Kang, Taegwan
Lee, Minwoo
Chae, Younhyung
Jung, Kyomin
contents While large language models (LLMs) are increasingly used as automatic judges for question answering (QA) and other reference-conditioned evaluation tasks, little is known about their ability to adhere to a provided reference. We identify a critical failure mode of such reference-based LLM QA evaluation: when the provided reference conflicts with the judge model's parametric knowledge, the resulting scores become unreliable, substantially degrading evaluation fidelity. To study this phenomenon systematically, we introduce a controlled swapped-reference QA framework that induces reference-belief conflicts. Specifically, we replace the reference answer with an incorrect entity and construct diverse pairings of original and swapped references with correspondingly aligned candidate answers. Surprisingly, grading reliability drops sharply under swapped references across a broad set of judge models. We empirically show that this vulnerability is driven by judges' over-reliance on parametric knowledge, leading judges to disregard the given reference under conflict. Finally, we find that this failure persists under common prompt-based mitigation strategies, highlighting a fundamental limitation of LLM-as-a-judge evaluation and motivating reference-based protocols that enforce stronger adherence to the provided reference.
format Preprint
id arxiv_https___arxiv_org_abs_2601_07506
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
Lee, Dongryeol
Hwang, Yerin
Kang, Taegwan
Lee, Minwoo
Chae, Younhyung
Jung, Kyomin
Computation and Language
While large language models (LLMs) are increasingly used as automatic judges for question answering (QA) and other reference-conditioned evaluation tasks, little is known about their ability to adhere to a provided reference. We identify a critical failure mode of such reference-based LLM QA evaluation: when the provided reference conflicts with the judge model's parametric knowledge, the resulting scores become unreliable, substantially degrading evaluation fidelity. To study this phenomenon systematically, we introduce a controlled swapped-reference QA framework that induces reference-belief conflicts. Specifically, we replace the reference answer with an incorrect entity and construct diverse pairings of original and swapped references with correspondingly aligned candidate answers. Surprisingly, grading reliability drops sharply under swapped references across a broad set of judge models. We empirically show that this vulnerability is driven by judges' over-reliance on parametric knowledge, leading judges to disregard the given reference under conflict. Finally, we find that this failure persists under common prompt-based mitigation strategies, highlighting a fundamental limitation of LLM-as-a-judge evaluation and motivating reference-based protocols that enforce stronger adherence to the provided reference.
title Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
topic Computation and Language
url https://arxiv.org/abs/2601.07506