Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Jiangnan, Liu, Cheng-Tse, Deilamsalehy, Hanieh, Ahmed, Nesreen K., Mathur, Puneet, Lipka, Nedim, Dernoncourt, Franck, Rossi, Ryan A.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917257241690112
author Fang, Jiangnan
Liu, Cheng-Tse
Deilamsalehy, Hanieh
Ahmed, Nesreen K.
Mathur, Puneet
Lipka, Nedim
Dernoncourt, Franck
Rossi, Ryan A.
author_facet Fang, Jiangnan
Liu, Cheng-Tse
Deilamsalehy, Hanieh
Ahmed, Nesreen K.
Mathur, Puneet
Lipka, Nedim
Dernoncourt, Franck
Rossi, Ryan A.
contents Large language model (LLM) judges have often been used alongside traditional, algorithm-based metrics for tasks like summarization because they better capture semantic information, are better at reasoning, and are more robust to paraphrasing. However, LLM judges show biases for length and order among others, and are vulnerable to various adversarial input prompts. While recent studies have looked into these biases, few have analyzed them at a more granular level in relation to a well-defined overlap metric. In this work we provide an LLM judge bias analysis as a function of overlap with human-written responses in the domain of summarization. We test 9 recent LLMs with parameter counts ranging from 1 billion to 12 billion, including variants of Gemma 3 and LLaMA 3. We find that LLM judges increasingly prefer summaries generated by other LLMs over those written by humans as the similarities (as measured by ROUGE and BLEU) between the judged summaries decrease, and this pattern extends to all but one model tested, and exists regardless of the models' own position biases. Additionally, we find that models struggle to judge even summaries with limited overlaps, suggesting that LLM-as-a-judge in the summary domain should rely on techniques beyond a simple comparison.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07673
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation
Fang, Jiangnan
Liu, Cheng-Tse
Deilamsalehy, Hanieh
Ahmed, Nesreen K.
Mathur, Puneet
Lipka, Nedim
Dernoncourt, Franck
Rossi, Ryan A.
Computation and Language
Large language model (LLM) judges have often been used alongside traditional, algorithm-based metrics for tasks like summarization because they better capture semantic information, are better at reasoning, and are more robust to paraphrasing. However, LLM judges show biases for length and order among others, and are vulnerable to various adversarial input prompts. While recent studies have looked into these biases, few have analyzed them at a more granular level in relation to a well-defined overlap metric. In this work we provide an LLM judge bias analysis as a function of overlap with human-written responses in the domain of summarization. We test 9 recent LLMs with parameter counts ranging from 1 billion to 12 billion, including variants of Gemma 3 and LLaMA 3. We find that LLM judges increasingly prefer summaries generated by other LLMs over those written by humans as the similarities (as measured by ROUGE and BLEU) between the judged summaries decrease, and this pattern extends to all but one model tested, and exists regardless of the models' own position biases. Additionally, we find that models struggle to judge even summaries with limited overlaps, suggesting that LLM-as-a-judge in the summary domain should rely on techniques beyond a simple comparison.
title Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation
topic Computation and Language
url https://arxiv.org/abs/2602.07673