CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Zhiyuan, Li, Chenliang, Shi, Yingcheng, Shen, Weizhou, Yan, Ming, Huang, Fei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915957212971008
author Lu, Zhiyuan
Li, Chenliang
Shi, Yingcheng
Shen, Weizhou
Yan, Ming
Huang, Fei
author_facet Lu, Zhiyuan
Li, Chenliang
Shi, Yingcheng
Shen, Weizhou
Yan, Ming
Huang, Fei
contents While large language models now handle million-token contexts, their capacity for reasoning across entire document repositories remains largely untested. Existing benchmarks are inadequate, as they are mostly limited to single long texts or rely on a "sparse retrieval" assumption-that answers can be derived from a few relevant chunks. This assumption fails for true corpus-level analysis, where evidence is highly dispersed across hundreds of documents and answers require global integration, comparison, and statistical aggregation. To address this critical gap, we introduce CorpusQA, a new benchmark scaling up to 10 million tokens, generated via a novel data synthesis framework. By decoupling reasoning from textual representation, this framework creates complex, computation-intensive queries with programmatically guaranteed ground-truth answers, challenging systems to perform holistic reasoning over vast, unstructured text without relying on fallible human annotation. We further demonstrate the utility of our framework beyond evaluation, showing that fine-tuning on our synthesized data effectively enhances an LLM's general long-context reasoning capabilities. Extensive experiments reveal that even state-of-the-art long-context LLMs struggle as input length increases, and standard retrieval-augmented generation systems collapse entirely. Our findings indicate that memory-augmented agentic architectures offer a more robust alternative, suggesting a critical shift is needed from simply extending context windows to developing advanced architectures for global information synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14952
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
Lu, Zhiyuan
Li, Chenliang
Shi, Yingcheng
Shen, Weizhou
Yan, Ming
Huang, Fei
Computation and Language
Artificial Intelligence
While large language models now handle million-token contexts, their capacity for reasoning across entire document repositories remains largely untested. Existing benchmarks are inadequate, as they are mostly limited to single long texts or rely on a "sparse retrieval" assumption-that answers can be derived from a few relevant chunks. This assumption fails for true corpus-level analysis, where evidence is highly dispersed across hundreds of documents and answers require global integration, comparison, and statistical aggregation. To address this critical gap, we introduce CorpusQA, a new benchmark scaling up to 10 million tokens, generated via a novel data synthesis framework. By decoupling reasoning from textual representation, this framework creates complex, computation-intensive queries with programmatically guaranteed ground-truth answers, challenging systems to perform holistic reasoning over vast, unstructured text without relying on fallible human annotation. We further demonstrate the utility of our framework beyond evaluation, showing that fine-tuning on our synthesized data effectively enhances an LLM's general long-context reasoning capabilities. Extensive experiments reveal that even state-of-the-art long-context LLMs struggle as input length increases, and standard retrieval-augmented generation systems collapse entirely. Our findings indicate that memory-augmented agentic architectures offer a more robust alternative, suggesting a critical shift is needed from simply extending context windows to developing advanced architectures for global information synthesis.
title CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.14952