Document Reconstruction Unlocks Scalable Long-Context RLVR

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xiao, Yao, Wang, Lei, Deng, Yue, Chen, Guanzheng, Jin, Ziqi, Kim, Jung-jae, Li, Xiaoli, Lee, Roy Ka-wei, Bing, Lidong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912936653488128
author Xiao, Yao
Wang, Lei
Deng, Yue
Chen, Guanzheng
Jin, Ziqi
Kim, Jung-jae
Li, Xiaoli
Lee, Roy Ka-wei
Bing, Lidong
author_facet Xiao, Yao
Wang, Lei
Deng, Yue
Chen, Guanzheng
Jin, Ziqi
Kim, Jung-jae
Li, Xiaoli
Lee, Roy Ka-wei
Bing, Lidong
contents Reinforcement Learning with Verifiable Rewards~(RLVR) has become a prominent paradigm to enhance the capabilities (i.e.\ long-context) of Large Language Models~(LLMs). However, it often relies on gold-standard answers or explicit evaluation rubrics provided by powerful teacher models or human experts, which are costly and time-consuming. In this work, we investigate unsupervised approaches to enhance the long-context capabilities of LLMs, eliminating the need for heavy human annotations or teacher models' supervision. Specifically, we first replace a few paragraphs with special placeholders in a long document. LLMs are trained through reinforcement learning to reconstruct the document by correctly identifying and sequencing missing paragraphs from a set of candidate options. This training paradigm enables the model to capture global narrative coherence, significantly boosting long-context performance. We validate the effectiveness of our method on two widely used benchmarks, RULER and LongBench~v2. While acquiring noticeable gains on RULER, it can also achieve a reasonable improvement on LongBench~v2 without any manually curated long-context QA data. Furthermore, we conduct extensive ablation studies to analyze the impact of reward design, data curation strategies, training schemes, and data scaling effects on model performance. We publicly release our code, data, and models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08237
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Document Reconstruction Unlocks Scalable Long-Context RLVR
Xiao, Yao
Wang, Lei
Deng, Yue
Chen, Guanzheng
Jin, Ziqi
Kim, Jung-jae
Li, Xiaoli
Lee, Roy Ka-wei
Bing, Lidong
Computation and Language
Reinforcement Learning with Verifiable Rewards~(RLVR) has become a prominent paradigm to enhance the capabilities (i.e.\ long-context) of Large Language Models~(LLMs). However, it often relies on gold-standard answers or explicit evaluation rubrics provided by powerful teacher models or human experts, which are costly and time-consuming. In this work, we investigate unsupervised approaches to enhance the long-context capabilities of LLMs, eliminating the need for heavy human annotations or teacher models' supervision. Specifically, we first replace a few paragraphs with special placeholders in a long document. LLMs are trained through reinforcement learning to reconstruct the document by correctly identifying and sequencing missing paragraphs from a set of candidate options. This training paradigm enables the model to capture global narrative coherence, significantly boosting long-context performance. We validate the effectiveness of our method on two widely used benchmarks, RULER and LongBench~v2. While acquiring noticeable gains on RULER, it can also achieve a reasonable improvement on LongBench~v2 without any manually curated long-context QA data. Furthermore, we conduct extensive ablation studies to analyze the impact of reward design, data curation strategies, training schemes, and data scaling effects on model performance. We publicly release our code, data, and models.
title Document Reconstruction Unlocks Scalable Long-Context RLVR
topic Computation and Language
url https://arxiv.org/abs/2602.08237