Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Futing, Yan, Jianhao, Luo, Yun, Cui, Ganqu, Wang, Zhi, Qu, Xiaoye, Zhang, Yue, Cheng, Yu, Lin, Tao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914324803485696
author Wang, Futing
Yan, Jianhao
Luo, Yun
Cui, Ganqu
Wang, Zhi
Qu, Xiaoye
Zhang, Yue
Cheng, Yu
Lin, Tao
author_facet Wang, Futing
Yan, Jianhao
Luo, Yun
Cui, Ganqu
Wang, Zhi
Qu, Xiaoye
Zhang, Yue
Cheng, Yu
Lin, Tao
contents Achieving effective test-time scaling requires models to engage in In-Context Exploration -- the intrinsic ability to generate, verify, and refine multiple reasoning hypotheses within a single continuous context. Grounded in State Coverage theory, our analysis identifies a critical bottleneck to enabling this capability: while broader state coverage requires longer reasoning trajectories, the probability of sampling such sequences decays exponentially during autoregressive generation, a phenomenon we term the ``Shallow Exploration Trap''. To bridge this gap, we propose Length-Incentivized Exploration(\method). This simple yet effective recipe explicitly encourages models to explore more via a length-based reward coupled with a redundancy penalty, thereby maximizing state coverage in two-step manner. Comprehensive experiments across different models (Qwen3, Llama) demonstrate that \method effectively incentivize in-context exploration. As a result, our method achieves an average improvement of 4.4\% on in-domain tasks and a 2.7\% gain on out-of-domain benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_11748
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning
Wang, Futing
Yan, Jianhao
Luo, Yun
Cui, Ganqu
Wang, Zhi
Qu, Xiaoye
Zhang, Yue
Cheng, Yu
Lin, Tao
Computation and Language
Achieving effective test-time scaling requires models to engage in In-Context Exploration -- the intrinsic ability to generate, verify, and refine multiple reasoning hypotheses within a single continuous context. Grounded in State Coverage theory, our analysis identifies a critical bottleneck to enabling this capability: while broader state coverage requires longer reasoning trajectories, the probability of sampling such sequences decays exponentially during autoregressive generation, a phenomenon we term the ``Shallow Exploration Trap''. To bridge this gap, we propose Length-Incentivized Exploration(\method). This simple yet effective recipe explicitly encourages models to explore more via a length-based reward coupled with a redundancy penalty, thereby maximizing state coverage in two-step manner. Comprehensive experiments across different models (Qwen3, Llama) demonstrate that \method effectively incentivize in-context exploration. As a result, our method achieves an average improvement of 4.4\% on in-domain tasks and a 2.7\% gain on out-of-domain benchmarks.
title Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning
topic Computation and Language
url https://arxiv.org/abs/2602.11748