Behavioral Analysis of Information Salience in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Trienes, Jan, Schlötterer, Jörg, Li, Junyi Jessy, Seifert, Christin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908381066821632
author Trienes, Jan
Schlötterer, Jörg
Li, Junyi Jessy
Seifert, Christin
author_facet Trienes, Jan
Schlötterer, Jörg
Li, Junyi Jessy
Seifert, Christin
contents Large Language Models (LLMs) excel at text summarization, a task that requires models to select content based on its importance. However, the exact notion of salience that LLMs have internalized remains unclear. To bridge this gap, we introduce an explainable framework to systematically derive and investigate information salience in LLMs through their summarization behavior. Using length-controlled summarization as a behavioral probe into the content selection process, and tracing the answerability of Questions Under Discussion throughout, we derive a proxy for how models prioritize information. Our experiments on 13 models across four datasets reveal that LLMs have a nuanced, hierarchical notion of salience, generally consistent across model families and sizes. While models show highly consistent behavior and hence salience patterns, this notion of salience cannot be accessed through introspection, and only weakly correlates with human perceptions of information salience.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14613
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Behavioral Analysis of Information Salience in Large Language Models
Trienes, Jan
Schlötterer, Jörg
Li, Junyi Jessy
Seifert, Christin
Computation and Language
Large Language Models (LLMs) excel at text summarization, a task that requires models to select content based on its importance. However, the exact notion of salience that LLMs have internalized remains unclear. To bridge this gap, we introduce an explainable framework to systematically derive and investigate information salience in LLMs through their summarization behavior. Using length-controlled summarization as a behavioral probe into the content selection process, and tracing the answerability of Questions Under Discussion throughout, we derive a proxy for how models prioritize information. Our experiments on 13 models across four datasets reveal that LLMs have a nuanced, hierarchical notion of salience, generally consistent across model families and sizes. While models show highly consistent behavior and hence salience patterns, this notion of salience cannot be accessed through introspection, and only weakly correlates with human perceptions of information salience.
title Behavioral Analysis of Information Salience in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2502.14613