The simulation of judgment in LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Loru, Edoardo, Nudo, Jacopo, Di Marco, Niccolò, Santirocchi, Alessandro, Atzeni, Roberto, Cinelli, Matteo, Cestari, Vincenzo, Rossi-Arnaud, Clelia, Quattrociocchi, Walter
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912652612075520
author Loru, Edoardo
Nudo, Jacopo
Di Marco, Niccolò
Santirocchi, Alessandro
Atzeni, Roberto
Cinelli, Matteo
Cestari, Vincenzo
Rossi-Arnaud, Clelia
Quattrociocchi, Walter
author_facet Loru, Edoardo
Nudo, Jacopo
Di Marco, Niccolò
Santirocchi, Alessandro
Atzeni, Roberto
Cinelli, Matteo
Cestari, Vincenzo
Rossi-Arnaud, Clelia
Quattrociocchi, Walter
contents Large Language Models (LLMs) are increasingly embedded in evaluative processes, from information filtering to assessing and addressing knowledge gaps through explanation and credibility judgments. This raises the need to examine how such evaluations are built, what assumptions they rely on, and how their strategies diverge from those of humans. We benchmark six LLMs against expert ratings--NewsGuard and Media Bias/Fact Check--and against human judgments collected through a controlled experiment. We use news domains purely as a controlled benchmark for evaluative tasks, focusing on the underlying mechanisms rather than on news classification per se. To enable direct comparison, we implement a structured agentic framework in which both models and nonexpert participants follow the same evaluation procedure: selecting criteria, retrieving content, and producing justifications. Despite output alignment, our findings show consistent differences in the observable criteria guiding model evaluations, suggesting that lexical associations and statistical priors could influence evaluations in ways that differ from contextual reasoning. This reliance is associated with systematic effects: political asymmetries and a tendency to confuse linguistic form with epistemic reliability--a dynamic we term epistemia, the illusion of knowledge that emerges when surface plausibility replaces verification. Indeed, delegating judgment to such systems may affect the heuristics underlying evaluative processes, suggesting a shift from normative reasoning toward pattern-based approximation and raising open questions about the role of LLMs in evaluative processes.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04426
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The simulation of judgment in LLMs
Loru, Edoardo
Nudo, Jacopo
Di Marco, Niccolò
Santirocchi, Alessandro
Atzeni, Roberto
Cinelli, Matteo
Cestari, Vincenzo
Rossi-Arnaud, Clelia
Quattrociocchi, Walter
Computation and Language
Artificial Intelligence
Computers and Society
Large Language Models (LLMs) are increasingly embedded in evaluative processes, from information filtering to assessing and addressing knowledge gaps through explanation and credibility judgments. This raises the need to examine how such evaluations are built, what assumptions they rely on, and how their strategies diverge from those of humans. We benchmark six LLMs against expert ratings--NewsGuard and Media Bias/Fact Check--and against human judgments collected through a controlled experiment. We use news domains purely as a controlled benchmark for evaluative tasks, focusing on the underlying mechanisms rather than on news classification per se. To enable direct comparison, we implement a structured agentic framework in which both models and nonexpert participants follow the same evaluation procedure: selecting criteria, retrieving content, and producing justifications. Despite output alignment, our findings show consistent differences in the observable criteria guiding model evaluations, suggesting that lexical associations and statistical priors could influence evaluations in ways that differ from contextual reasoning. This reliance is associated with systematic effects: political asymmetries and a tendency to confuse linguistic form with epistemic reliability--a dynamic we term epistemia, the illusion of knowledge that emerges when surface plausibility replaces verification. Indeed, delegating judgment to such systems may affect the heuristics underlying evaluative processes, suggesting a shift from normative reasoning toward pattern-based approximation and raising open questions about the role of LLMs in evaluative processes.
title The simulation of judgment in LLMs
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2502.04426