Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sheng, Boheng, Yao, Jiacheng, Zhang, Meicong, He, Guoxiu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915319737483264
author Sheng, Boheng
Yao, Jiacheng
Zhang, Meicong
He, Guoxiu
author_facet Sheng, Boheng
Yao, Jiacheng
Zhang, Meicong
He, Guoxiu
contents Large language models (LLMs) often struggle to accurately read and comprehend extremely long texts. Current methods for improvement typically rely on splitting long contexts into fixed-length chunks. However, fixed truncation risks separating semantically relevant content, leading to ambiguity and compromising accurate understanding. To overcome this limitation, we propose a straightforward approach for dynamically separating and selecting chunks of long context, facilitating a more streamlined input for LLMs. In particular, we compute semantic similarities between adjacent sentences, using lower similarities to adaptively divide long contexts into variable-length chunks. We further train a question-aware classifier to select sensitive chunks that are critical for answering specific questions. Experimental results on both single-hop and multi-hop question-answering benchmarks show that the proposed approach consistently outperforms strong baselines. Notably, it maintains robustness across a wide range of input lengths, handling sequences of up to 256k tokens. Our datasets and code are available at the following link: https://github.com/ECNU-Text-Computing/DCS
format Preprint
id arxiv_https___arxiv_org_abs_2506_00773
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language Models
Sheng, Boheng
Yao, Jiacheng
Zhang, Meicong
He, Guoxiu
Computation and Language
Large language models (LLMs) often struggle to accurately read and comprehend extremely long texts. Current methods for improvement typically rely on splitting long contexts into fixed-length chunks. However, fixed truncation risks separating semantically relevant content, leading to ambiguity and compromising accurate understanding. To overcome this limitation, we propose a straightforward approach for dynamically separating and selecting chunks of long context, facilitating a more streamlined input for LLMs. In particular, we compute semantic similarities between adjacent sentences, using lower similarities to adaptively divide long contexts into variable-length chunks. We further train a question-aware classifier to select sensitive chunks that are critical for answering specific questions. Experimental results on both single-hop and multi-hop question-answering benchmarks show that the proposed approach consistently outperforms strong baselines. Notably, it maintains robustness across a wide range of input lengths, handling sequences of up to 256k tokens. Our datasets and code are available at the following link: https://github.com/ECNU-Text-Computing/DCS
title Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2506.00773