Hallucinate at the Last in Long Response Generation: A Case Study on Long Document Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Joonho, Yoon, Seunghyun, Chang, Hwan, Kim, Byeongjeong, Lee, Hwanhee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912819115458560
author Yang, Joonho
Yoon, Seunghyun
Chang, Hwan
Kim, Byeongjeong
Lee, Hwanhee
author_facet Yang, Joonho
Yoon, Seunghyun
Chang, Hwan
Kim, Byeongjeong
Lee, Hwanhee
contents Large Language Models (LLMs) have significantly advanced text generation capabilities, including tasks like summarization, often producing coherent and fluent outputs. However, faithfulness to source material remains a significant challenge due to the generation of hallucinations. While extensive research focuses on detecting and reducing these inaccuracies, less attention has been paid to the positional distribution of hallucination within generated text, particularly in long outputs. In this work, we investigate where hallucinations occur in LLM-based long response generation, using long document summarization as a key case study. Focusing on the challenging setting of long context-aware long response generation, we find a consistent and concerning phenomenon: hallucinations tend to concentrate disproportionately in the latter parts of the generated long response. To understand this bias, we explore potential contributing factors related to the dynamics of attention and decoding over long sequences. Furthermore, we investigate methods to mitigate this positional hallucination, aiming to improve faithfulness specifically in the concluding segments of long outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15291
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hallucinate at the Last in Long Response Generation: A Case Study on Long Document Summarization
Yang, Joonho
Yoon, Seunghyun
Chang, Hwan
Kim, Byeongjeong
Lee, Hwanhee
Computation and Language
Large Language Models (LLMs) have significantly advanced text generation capabilities, including tasks like summarization, often producing coherent and fluent outputs. However, faithfulness to source material remains a significant challenge due to the generation of hallucinations. While extensive research focuses on detecting and reducing these inaccuracies, less attention has been paid to the positional distribution of hallucination within generated text, particularly in long outputs. In this work, we investigate where hallucinations occur in LLM-based long response generation, using long document summarization as a key case study. Focusing on the challenging setting of long context-aware long response generation, we find a consistent and concerning phenomenon: hallucinations tend to concentrate disproportionately in the latter parts of the generated long response. To understand this bias, we explore potential contributing factors related to the dynamics of attention and decoding over long sequences. Furthermore, we investigate methods to mitigate this positional hallucination, aiming to improve faithfulness specifically in the concluding segments of long outputs.
title Hallucinate at the Last in Long Response Generation: A Case Study on Long Document Summarization
topic Computation and Language
url https://arxiv.org/abs/2505.15291