LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dietz, Laura, Zendel, Oleg, Bailey, Peter, Clarke, Charles, Cotterill, Ellese, Dalton, Jeff, Hasibi, Faegheh, Sanderson, Mark, Craswell, Nick
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917210593689600
author Dietz, Laura
Zendel, Oleg
Bailey, Peter
Clarke, Charles
Cotterill, Ellese
Dalton, Jeff
Hasibi, Faegheh
Sanderson, Mark
Craswell, Nick
author_facet Dietz, Laura
Zendel, Oleg
Bailey, Peter
Clarke, Charles
Cotterill, Ellese
Dalton, Jeff
Hasibi, Faegheh
Sanderson, Mark
Craswell, Nick
contents Large Language Models (LLMs) are increasingly used to evaluate information retrieval (IR) systems, generating relevance judgments traditionally made by human assessors. Recent empirical studies suggest that LLM-based evaluations often align with human judgments, leading some to suggest that human judges may no longer be necessary, while others highlight concerns about judgment reliability, validity, and long-term impact. As IR systems begin incorporating LLM-generated signals, evaluation outcomes risk becoming self-reinforcing, potentially leading to misleading conclusions. This paper examines scenarios where LLM-evaluators may falsely indicate success, particularly when LLM-based judgments influence both system development and evaluation. We highlight key risks, including bias reinforcement, reproducibility challenges, and inconsistencies in assessment methodologies. To address these concerns, we propose tests to quantify adverse effects, guardrails, and a collaborative framework for constructing reusable test collections that integrate LLM judgments responsibly. By providing perspectives from academia and industry, this work aims to establish best practices for the principled use of LLMs in IR evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19076
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
Dietz, Laura
Zendel, Oleg
Bailey, Peter
Clarke, Charles
Cotterill, Ellese
Dalton, Jeff
Hasibi, Faegheh
Sanderson, Mark
Craswell, Nick
Information Retrieval
H.3.3
Large Language Models (LLMs) are increasingly used to evaluate information retrieval (IR) systems, generating relevance judgments traditionally made by human assessors. Recent empirical studies suggest that LLM-based evaluations often align with human judgments, leading some to suggest that human judges may no longer be necessary, while others highlight concerns about judgment reliability, validity, and long-term impact. As IR systems begin incorporating LLM-generated signals, evaluation outcomes risk becoming self-reinforcing, potentially leading to misleading conclusions. This paper examines scenarios where LLM-evaluators may falsely indicate success, particularly when LLM-based judgments influence both system development and evaluation. We highlight key risks, including bias reinforcement, reproducibility challenges, and inconsistencies in assessment methodologies. To address these concerns, we propose tests to quantify adverse effects, guardrails, and a collaborative framework for constructing reusable test collections that integrate LLM judgments responsibly. By providing perspectives from academia and industry, this work aims to establish best practices for the principled use of LLMs in IR evaluation.
title LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
topic Information Retrieval
H.3.3
url https://arxiv.org/abs/2504.19076