Document Reconstruction Unlocks Scalable Long-Context RLVR
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Yao, Wang, Lei, Deng, Yue, Chen, Guanzheng, Jin, Ziqi, Kim, Jung-jae, Li, Xiaoli, Lee, Roy Ka-wei, Bing, Lidong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting Self-Play Preference Optimization: On the Role of Prompt Difficulty
by: Xiao, Yao, et al.
Published: (2025)
by: Xiao, Yao, et al.
Published: (2025)
LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards
by: Chen, Guanzheng, et al.
Published: (2026)
by: Chen, Guanzheng, et al.
Published: (2026)
LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization
by: Chen, Guanzheng, et al.
Published: (2025)
by: Chen, Guanzheng, et al.
Published: (2025)
Finding the Sweet Spot: Preference Data Construction for Scaling Preference Optimization
by: Xiao, Yao, et al.
Published: (2025)
by: Xiao, Yao, et al.
Published: (2025)
MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
by: Yang, Zonglin, et al.
Published: (2026)
by: Yang, Zonglin, et al.
Published: (2026)
CLEX: Continuous Length Extrapolation for Large Language Models
by: Chen, Guanzheng, et al.
Published: (2023)
by: Chen, Guanzheng, et al.
Published: (2023)
Easy-to-Implement Two-Way Effect Decomposition for Any Outcome Variable with Endogenous Mediator
by: Kim, Bora, et al.
Published: (2025)
by: Kim, Bora, et al.
Published: (2025)
Finding network effect of randomized treatment under weak assumptions for any outcome and any effect heterogeneity
by: Lee, Myoung-jae
Published: (2025)
by: Lee, Myoung-jae
Published: (2025)
First Try Matters: Revisiting the Role of Reflection in Reasoning Models
by: Kang, Liwei, et al.
Published: (2025)
by: Kang, Liwei, et al.
Published: (2025)
Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models
by: Luo, Ziwei, et al.
Published: (2026)
by: Luo, Ziwei, et al.
Published: (2026)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations
by: Xiao, Yunze, et al.
Published: (2024)
by: Xiao, Yunze, et al.
Published: (2024)
On the Role of Discreteness in Diffusion LLMs
by: Jin, Ziqi, et al.
Published: (2025)
by: Jin, Ziqi, et al.
Published: (2025)
HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution
by: Hong, Hanhua, et al.
Published: (2026)
by: Hong, Hanhua, et al.
Published: (2026)
RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding
by: Chen, Guanzheng, et al.
Published: (2025)
by: Chen, Guanzheng, et al.
Published: (2025)
Evolving Prompts In-Context: An Open-ended, Self-replicating Perspective
by: Wang, Jianyu, et al.
Published: (2025)
by: Wang, Jianyu, et al.
Published: (2025)
Rethinking Multiple-Choice Questions for RLVR: Unlocking Potential via Distractor Design
by: Guo, Xu, et al.
Published: (2026)
by: Guo, Xu, et al.
Published: (2026)
Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
by: Wei, Chengwei, et al.
Published: (2025)
by: Wei, Chengwei, et al.
Published: (2025)
Multilingual Jailbreak Challenges in Large Language Models
by: Deng, Yue, et al.
Published: (2023)
by: Deng, Yue, et al.
Published: (2023)
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
by: Lee, Seung-jae, et al.
Published: (2025)
by: Lee, Seung-jae, et al.
Published: (2025)
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
by: Xie, Can, et al.
Published: (2025)
by: Xie, Can, et al.
Published: (2025)
TeXBLEU: Automatic Metric for Evaluate LaTeX Format
by: Jung, Kyudan, et al.
Published: (2024)
by: Jung, Kyudan, et al.
Published: (2024)
Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
by: Zhang, Yuheng, et al.
Published: (2025)
by: Zhang, Yuheng, et al.
Published: (2025)
Shifting Long-Context LLMs Research from Input to Output
by: Wu, Yuhao, et al.
Published: (2025)
by: Wu, Yuhao, et al.
Published: (2025)
ParaICL: Towards Parallel In-Context Learning
by: Li, Xingxuan, et al.
Published: (2024)
by: Li, Xingxuan, et al.
Published: (2024)
Optimizing OLED performance on polyimide substrates: Evaluation of ITO and organic layer thicknesses with different encapsulation materials
by: Hyunsu Cho, et al.
Published: (2025)
by: Hyunsu Cho, et al.
Published: (2025)
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
by: Yang, Zhicheng, et al.
Published: (2025)
by: Yang, Zhicheng, et al.
Published: (2025)
PACIT: Unlocking the Power of Examples for Better In-Context Instruction Tuning
by: Xue, Tianci, et al.
Published: (2023)
by: Xue, Tianci, et al.
Published: (2023)
CoinMath: Harnessing the Power of Coding Instruction for Math LLMs
by: Wei, Chengwei, et al.
Published: (2024)
by: Wei, Chengwei, et al.
Published: (2024)
InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning
by: Wei, Chengwei, et al.
Published: (2026)
by: Wei, Chengwei, et al.
Published: (2026)
Self-Distilled RLVR
by: Yang, Chenxu, et al.
Published: (2026)
by: Yang, Chenxu, et al.
Published: (2026)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
by: Shi, Wenhao, et al.
Published: (2024)
by: Shi, Wenhao, et al.
Published: (2024)
Unlocking Temporal Question Answering for Large Language Models with Tailor-Made Reasoning Logic
by: Li, Xingxuan, et al.
Published: (2023)
by: Li, Xingxuan, et al.
Published: (2023)
Towards Scalability and Extensibility of Query Reformulation Modeling in E-commerce Search
by: Zhang, Ziqi, et al.
Published: (2024)
by: Zhang, Ziqi, et al.
Published: (2024)
OPSD Compresses What RLVR Teaches: A Post-RL Compaction Stage for Reasoning Models
by: Kim, Jaehoon, et al.
Published: (2026)
by: Kim, Jaehoon, et al.
Published: (2026)
Detecting RLVR Training Data via Structural Convergence of Reasoning
by: Zhang, Hongbo, et al.
Published: (2026)
by: Zhang, Hongbo, et al.
Published: (2026)
Bridging Modalities: Enhancing Cross-Modality Hate Speech Detection with Few-Shot In-Context Learning
by: Hee, Ming Shan, et al.
Published: (2024)
by: Hee, Ming Shan, et al.
Published: (2024)
Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets
by: Joshi, Harshit, et al.
Published: (2026)
by: Joshi, Harshit, et al.
Published: (2026)
LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
by: Liu, Xiaoran, et al.
Published: (2025)
by: Liu, Xiaoran, et al.
Published: (2025)
Similar Items
-
Revisiting Self-Play Preference Optimization: On the Role of Prompt Difficulty
by: Xiao, Yao, et al.
Published: (2025) -
LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards
by: Chen, Guanzheng, et al.
Published: (2026) -
LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization
by: Chen, Guanzheng, et al.
Published: (2025) -
Finding the Sweet Spot: Preference Data Construction for Scaling Preference Optimization
by: Xiao, Yao, et al.
Published: (2025) -
MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
by: Yang, Zonglin, et al.
Published: (2026)