Lost in Stories: Consistency Bugs in Long Story Generation by LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Junjie, Guo, Xinrui, Wu, Yuhao, Lee, Roy Ka-Wei, Li, Hongzhi, Xie, Yutao
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914374652788736
author Li, Junjie
Guo, Xinrui
Wu, Yuhao
Lee, Roy Ka-Wei
Li, Hongzhi
Xie, Yutao
author_facet Li, Junjie
Guo, Xinrui
Wu, Yuhao
Lee, Roy Ka-Wei
Li, Hongzhi
Xie, Yutao
contents What happens when a storyteller forgets its own story? Large Language Models (LLMs) can now generate narratives spanning tens of thousands of words, but they often fail to maintain consistency throughout. When generating long-form narratives, these models can contradict their own established facts, character traits, and world rules. Existing story generation benchmarks focus mainly on plot quality and fluency, leaving consistency errors largely unexplored. To address this gap, we present ConStory-Bench, a benchmark designed to evaluate narrative consistency in long-form story generation. It contains 2,000 prompts across four task scenarios and defines a taxonomy of five error categories with 19 fine-grained subtypes. We also develop ConStory-Checker, an automated pipeline that detects contradictions and grounds each judgment in explicit textual evidence. Evaluating a range of LLMs through five research questions, we find that consistency errors show clear tendencies: they are most common in factual and temporal dimensions, tend to appear around the middle of narratives, occur in text segments with higher token-level entropy, and certain error types tend to co-occur. These findings can inform future efforts to improve consistency in long-form narrative generation. Our project page is available at https://picrew.github.io/constory-bench.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05890
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
Li, Junjie
Guo, Xinrui
Wu, Yuhao
Lee, Roy Ka-Wei
Li, Hongzhi
Xie, Yutao
Computation and Language
Artificial Intelligence
What happens when a storyteller forgets its own story? Large Language Models (LLMs) can now generate narratives spanning tens of thousands of words, but they often fail to maintain consistency throughout. When generating long-form narratives, these models can contradict their own established facts, character traits, and world rules. Existing story generation benchmarks focus mainly on plot quality and fluency, leaving consistency errors largely unexplored. To address this gap, we present ConStory-Bench, a benchmark designed to evaluate narrative consistency in long-form story generation. It contains 2,000 prompts across four task scenarios and defines a taxonomy of five error categories with 19 fine-grained subtypes. We also develop ConStory-Checker, an automated pipeline that detects contradictions and grounds each judgment in explicit textual evidence. Evaluating a range of LLMs through five research questions, we find that consistency errors show clear tendencies: they are most common in factual and temporal dimensions, tend to appear around the middle of narratives, occur in text segments with higher token-level entropy, and certain error types tend to co-occur. These findings can inform future efforts to improve consistency in long-form narrative generation. Our project page is available at https://picrew.github.io/constory-bench.github.io/.
title Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.05890