The Effect of Document Summarization on LLM-Based Relevance Judgments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mohtadi, Samaneh, Roitero, Kevin, Mizzaro, Stefano, Demartini, Gianluca
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915655954989056
author Mohtadi, Samaneh
Roitero, Kevin
Mizzaro, Stefano
Demartini, Gianluca
author_facet Mohtadi, Samaneh
Roitero, Kevin
Mizzaro, Stefano
Demartini, Gianluca
contents Relevance judgments are central to the evaluation of Information Retrieval (IR) systems, but obtaining them from human annotators is costly and time-consuming. Large Language Models (LLMs) have recently been proposed as automated assessors, showing promising alignment with human annotations. Most prior studies have treated documents as fixed units, feeding their full content directly to LLM assessors. We investigate how text summarization affects the reliability of LLM-based judgments and their downstream impact on IR evaluation. Using state-of-the-art LLMs across multiple TREC collections, we compare judgments made from full documents with those based on LLM-generated summaries of different lengths. We examine their agreement with human labels, their effect on retrieval effectiveness evaluation, and their influence on IR systems' ranking stability. Our findings show that summary-based judgments achieve comparable stability in systems' ranking to full-document judgments, while introducing systematic shifts in label distributions and biases that vary by model and dataset. These results highlight summarization as both an opportunity for more efficient large-scale IR evaluation and a methodological choice with important implications for the reliability of automatic judgments.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05334
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Effect of Document Summarization on LLM-Based Relevance Judgments
Mohtadi, Samaneh
Roitero, Kevin
Mizzaro, Stefano
Demartini, Gianluca
Information Retrieval
Artificial Intelligence
Computation and Language
Relevance judgments are central to the evaluation of Information Retrieval (IR) systems, but obtaining them from human annotators is costly and time-consuming. Large Language Models (LLMs) have recently been proposed as automated assessors, showing promising alignment with human annotations. Most prior studies have treated documents as fixed units, feeding their full content directly to LLM assessors. We investigate how text summarization affects the reliability of LLM-based judgments and their downstream impact on IR evaluation. Using state-of-the-art LLMs across multiple TREC collections, we compare judgments made from full documents with those based on LLM-generated summaries of different lengths. We examine their agreement with human labels, their effect on retrieval effectiveness evaluation, and their influence on IR systems' ranking stability. Our findings show that summary-based judgments achieve comparable stability in systems' ranking to full-document judgments, while introducing systematic shifts in label distributions and biases that vary by model and dataset. These results highlight summarization as both an opportunity for more efficient large-scale IR evaluation and a methodological choice with important implications for the reliability of automatic judgments.
title The Effect of Document Summarization on LLM-Based Relevance Judgments
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.05334