PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Mo, Chung, Tsz Ting, Zhou, Chulun, Li, Tong, Lu, Rui, Li, Jiangnan, Xu, Liyan, Lu, Haoshu, Zhang, Ning, Li, Jing, Zhou, Jie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909736492859392
author Yu, Mo
Chung, Tsz Ting
Zhou, Chulun
Li, Tong
Lu, Rui
Li, Jiangnan
Xu, Liyan
Lu, Haoshu
Zhang, Ning
Li, Jing
Zhou, Jie
author_facet Yu, Mo
Chung, Tsz Ting
Zhou, Chulun
Li, Tong
Lu, Rui
Li, Jiangnan
Xu, Liyan
Lu, Haoshu
Zhang, Ning
Li, Jing
Zhou, Jie
contents We introduce PRELUDE, a benchmark for evaluating long-context understanding through the task of determining whether a character's prequel story is consistent with the canonical narrative of the original book. Our task poses a stronger demand for global comprehension and deep reasoning than existing benchmarks -- as the prequels are not part of the original story, assessing their plausibility typically requires searching and integrating information that is only indirectly related. Empirically, 88% of instances require evidence from multiple parts of the narrative. Experimental results highlight the challenge of our task: in-context learning, RAG and in-domain training with state-of-the-art LLMs, and commercial DeepResearch services, lag behind humans by >15%. A further human study reveals that models often produce correct answers with flawed reasoning, leading to an over 30% gap in reasoning accuracy compared to humans. These findings underscore the substantial room for improvement in long-context understanding and reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09848
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts
Yu, Mo
Chung, Tsz Ting
Zhou, Chulun
Li, Tong
Lu, Rui
Li, Jiangnan
Xu, Liyan
Lu, Haoshu
Zhang, Ning
Li, Jing
Zhou, Jie
Computation and Language
Artificial Intelligence
We introduce PRELUDE, a benchmark for evaluating long-context understanding through the task of determining whether a character's prequel story is consistent with the canonical narrative of the original book. Our task poses a stronger demand for global comprehension and deep reasoning than existing benchmarks -- as the prequels are not part of the original story, assessing their plausibility typically requires searching and integrating information that is only indirectly related. Empirically, 88% of instances require evidence from multiple parts of the narrative. Experimental results highlight the challenge of our task: in-context learning, RAG and in-domain training with state-of-the-art LLMs, and commercial DeepResearch services, lag behind humans by >15%. A further human study reveals that models often produce correct answers with flawed reasoning, leading to an over 30% gap in reasoning accuracy compared to humans. These findings underscore the substantial room for improvement in long-context understanding and reasoning.
title PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.09848