CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Zhiyuan, Li, Chenliang, Shi, Yingcheng, Shen, Weizhou, Yan, Ming, Huang, Fei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Incentivizing In-depth Reasoning over Long Contexts with Process Advantage Shaping
by: Peng, Miao, et al.
Published: (2026)
by: Peng, Miao, et al.
Published: (2026)
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
by: Luo, Qi, et al.
Published: (2025)
by: Luo, Qi, et al.
Published: (2025)
Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
by: Shen, Weizhou, et al.
Published: (2024)
by: Shen, Weizhou, et al.
Published: (2024)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
by: Wang, Zengzhi, et al.
Published: (2023)
by: Wang, Zengzhi, et al.
Published: (2023)
Matina: A Large-Scale 73B Token Persian Text Corpus
by: Hosseinbeigi, Sara Bourbour, et al.
Published: (2025)
by: Hosseinbeigi, Sara Bourbour, et al.
Published: (2025)
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
by: Wan, Fanqi, et al.
Published: (2025)
by: Wan, Fanqi, et al.
Published: (2025)
LLMs and the Abstraction and Reasoning Corpus: Successes, Failures, and the Importance of Object-based Representations
by: Xu, Yudong, et al.
Published: (2023)
by: Xu, Yudong, et al.
Published: (2023)
SloPal: A 60-Million-Word Slovak Parliamentary Corpus with Aligned Speech and Fine-Tuned ASR Models
by: Božík, Erik, et al.
Published: (2025)
by: Božík, Erik, et al.
Published: (2025)
The PLLuM Instruction Corpus
by: Pęzik, Piotr, et al.
Published: (2025)
by: Pęzik, Piotr, et al.
Published: (2025)
Reasoning Abilities of Large Language Models: In-Depth Analysis on the Abstraction and Reasoning Corpus
by: Lee, Seungpil, et al.
Published: (2024)
by: Lee, Seungpil, et al.
Published: (2024)
NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
by: Silva, Enzo S. N., et al.
Published: (2026)
by: Silva, Enzo S. N., et al.
Published: (2026)
NESTLE: a No-Code Tool for Statistical Analysis of Legal Corpus
by: Cho, Kyoungyeon, et al.
Published: (2023)
by: Cho, Kyoungyeon, et al.
Published: (2023)
Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and Benchmark
by: Jiang, Feng, et al.
Published: (2023)
by: Jiang, Feng, et al.
Published: (2023)
Quantifying Geospatial in the Common Crawl Corpus
by: Ilyankou, Ilya, et al.
Published: (2024)
by: Ilyankou, Ilya, et al.
Published: (2024)
Tibyan Corpus: Balanced and Comprehensive Error Coverage Corpus Using ChatGPT for Arabic Grammatical Error Correction
by: Alrehili, Ahlam, et al.
Published: (2024)
by: Alrehili, Ahlam, et al.
Published: (2024)
LogicPrpBank: A Corpus for Logical Implication and Equivalence
by: Liu, Zhexiong, et al.
Published: (2024)
by: Liu, Zhexiong, et al.
Published: (2024)
PIIBench: A Unified Multi-Source Benchmark Corpus for Personally Identifiable Information Detection
by: Jha, Pritesh
Published: (2026)
by: Jha, Pritesh
Published: (2026)
OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages
by: Merx, Raphaël, et al.
Published: (2025)
by: Merx, Raphaël, et al.
Published: (2025)
Program Synthesis using Inductive Logic Programming for the Abstraction and Reasoning Corpus
by: Rocha, Filipe Marinho, et al.
Published: (2024)
by: Rocha, Filipe Marinho, et al.
Published: (2024)
QuranMorph: Morphologically Annotated Quranic Corpus
by: Akra, Diyam, et al.
Published: (2025)
by: Akra, Diyam, et al.
Published: (2025)
DRAGOn: Designing RAG On Periodically Updated Corpus
by: Chernogorskii, Fedor, et al.
Published: (2025)
by: Chernogorskii, Fedor, et al.
Published: (2025)
WritingBench: A Comprehensive Benchmark for Generative Writing
by: Wu, Yuning, et al.
Published: (2025)
by: Wu, Yuning, et al.
Published: (2025)
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
by: Huang, Tiancheng, et al.
Published: (2025)
by: Huang, Tiancheng, et al.
Published: (2025)
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
by: Tavakoli, Mohammad, et al.
Published: (2025)
by: Tavakoli, Mohammad, et al.
Published: (2025)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
by: Zuo, Yuxin, et al.
Published: (2025)
by: Zuo, Yuxin, et al.
Published: (2025)
QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Management
by: Shen, Weizhou, et al.
Published: (2025)
by: Shen, Weizhou, et al.
Published: (2025)
OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text
by: Li, Qingyun, et al.
Published: (2024)
by: Li, Qingyun, et al.
Published: (2024)
ACE-2005-PT: Corpus for Event Extraction in Portuguese
by: Cunha, Luís Filipe, et al.
Published: (2024)
by: Cunha, Luís Filipe, et al.
Published: (2024)
Corpus of Cross-lingual Dialogues with Minutes and Detection of Misunderstandings
by: Čechovič, Marko, et al.
Published: (2025)
by: Čechovič, Marko, et al.
Published: (2025)
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
by: Sharma, Rishikesh Kumar, et al.
Published: (2026)
HLDC: Hindi Legal Documents Corpus
by: Kapoor, Arnav, et al.
Published: (2022)
by: Kapoor, Arnav, et al.
Published: (2022)
NSINA: A News Corpus for Sinhala
by: Hettiarachchi, Hansi, et al.
Published: (2024)
by: Hettiarachchi, Hansi, et al.
Published: (2024)
R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning
by: Liu, Wanlong, et al.
Published: (2026)
by: Liu, Wanlong, et al.
Published: (2026)
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages using Wikidata
by: Sälevä, Jonne, et al.
Published: (2024)
by: Sälevä, Jonne, et al.
Published: (2024)
Integrating Linguistics and AI: Morphological Analysis and Corpus development of Endangered Toto Language of West Bengal
by: Guha, Ambalika, et al.
Published: (2025)
by: Guha, Ambalika, et al.
Published: (2025)
gaHealth: An English-Irish Bilingual Corpus of Health Data
by: Lankford, Séamus, et al.
Published: (2024)
by: Lankford, Séamus, et al.
Published: (2024)
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
DateLogicQA: Benchmarking Temporal Biases in Large Language Models
by: Bhatia, Gagan, et al.
Published: (2024)
by: Bhatia, Gagan, et al.
Published: (2024)
KPC-cF: Aspect-Based Sentiment Analysis via Implicit-Feature Alignment with Corpus Filtering
by: Nam, Kibeom
Published: (2024)
by: Nam, Kibeom
Published: (2024)
Refract ICL: Rethinking Example Selection in the Era of Million-Token Models
by: Akula, Arjun R., et al.
Published: (2025)
by: Akula, Arjun R., et al.
Published: (2025)
Similar Items
-
Incentivizing In-depth Reasoning over Long Contexts with Process Advantage Shaping
by: Peng, Miao, et al.
Published: (2026) -
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
by: Luo, Qi, et al.
Published: (2025) -
Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
by: Shen, Weizhou, et al.
Published: (2024) -
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
by: Wang, Zengzhi, et al.
Published: (2023) -
Matina: A Large-Scale 73B Token Persian Text Corpus
by: Hosseinbeigi, Sara Bourbour, et al.
Published: (2025)