Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914324803485696 |
|---|---|
| author | Wang, Futing Yan, Jianhao Luo, Yun Cui, Ganqu Wang, Zhi Qu, Xiaoye Zhang, Yue Cheng, Yu Lin, Tao |
| author_facet | Wang, Futing Yan, Jianhao Luo, Yun Cui, Ganqu Wang, Zhi Qu, Xiaoye Zhang, Yue Cheng, Yu Lin, Tao |
| contents | Achieving effective test-time scaling requires models to engage in In-Context Exploration -- the intrinsic ability to generate, verify, and refine multiple reasoning hypotheses within a single continuous context.
Grounded in State Coverage theory, our analysis identifies a critical bottleneck to enabling this capability: while broader state coverage requires longer reasoning trajectories, the probability of sampling such sequences decays exponentially during autoregressive generation, a phenomenon we term the ``Shallow Exploration Trap''.
To bridge this gap, we propose Length-Incentivized Exploration(\method).
This simple yet effective recipe explicitly encourages models to explore more via a length-based reward coupled with a redundancy penalty, thereby maximizing state coverage in two-step manner.
Comprehensive experiments across different models (Qwen3, Llama) demonstrate that \method effectively incentivize in-context exploration.
As a result, our method achieves an average improvement of 4.4\% on in-domain tasks and a 2.7\% gain on out-of-domain benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_11748 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning Wang, Futing Yan, Jianhao Luo, Yun Cui, Ganqu Wang, Zhi Qu, Xiaoye Zhang, Yue Cheng, Yu Lin, Tao Computation and Language Achieving effective test-time scaling requires models to engage in In-Context Exploration -- the intrinsic ability to generate, verify, and refine multiple reasoning hypotheses within a single continuous context. Grounded in State Coverage theory, our analysis identifies a critical bottleneck to enabling this capability: while broader state coverage requires longer reasoning trajectories, the probability of sampling such sequences decays exponentially during autoregressive generation, a phenomenon we term the ``Shallow Exploration Trap''. To bridge this gap, we propose Length-Incentivized Exploration(\method). This simple yet effective recipe explicitly encourages models to explore more via a length-based reward coupled with a redundancy penalty, thereby maximizing state coverage in two-step manner. Comprehensive experiments across different models (Qwen3, Llama) demonstrate that \method effectively incentivize in-context exploration. As a result, our method achieves an average improvement of 4.4\% on in-domain tasks and a 2.7\% gain on out-of-domain benchmarks. |
| title | Think Longer to Explore Deeper: Learn to Explore In-Context via Length-Incentivized Reinforcement Learning |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2602.11748 |