Saved in:
Bibliographic Details
Main Authors: Zou, Deyu, Chen, Yongqiang, Feng, Fan, Li, Mufei, Li, Pan, Gong, Yu, Cheng, James
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.12109
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916071546552320
author Zou, Deyu
Chen, Yongqiang
Feng, Fan
Li, Mufei
Li, Pan
Gong, Yu
Cheng, James
author_facet Zou, Deyu
Chen, Yongqiang
Feng, Fan
Li, Mufei
Li, Pan
Gong, Yu
Cheng, James
contents Reinforcement learning (RL) has become a de facto paradigm for building LLM-based agents that act, interact, and reason over extended task horizons. However, in active reasoning where agents must elicit new observations through interaction with the environment to solve the task, we find that outcome-based RL can induce a systematic failure mode which we call information self-locking (SeL): agents fail both to elicit informative feedback and to internalize obtained evidence. To understand the issue, we trace agentic behaviors into two coupled capabilities: Action Selection (AS), which determines observation streams, and Belief Tracking (BT), which updates the agent's internal task understanding. Theoretical and empirical analyses reveal a bidirectional bottleneck that leads to SeL: weak BT obscures the credit of informative actions, while weak AS deprives BT of useful evidence. This coupling weakens the learning signal for both capabilities and leads to SeL. To mitigate this issue, we propose AREW, a simple yet effective Advantage Reweighting method that uses easy-to-obtain directional critiques to reallocate credit within trajectories. Extensive experiments across 9 agentic tasks of varying complexity show that AREW significantly mitigates SeL, yielding up to 60-point gains in final performance. Code is available at https://github.com/unimpor/T3.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12109
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents
Zou, Deyu
Chen, Yongqiang
Feng, Fan
Li, Mufei
Li, Pan
Gong, Yu
Cheng, James
Artificial Intelligence
Reinforcement learning (RL) has become a de facto paradigm for building LLM-based agents that act, interact, and reason over extended task horizons. However, in active reasoning where agents must elicit new observations through interaction with the environment to solve the task, we find that outcome-based RL can induce a systematic failure mode which we call information self-locking (SeL): agents fail both to elicit informative feedback and to internalize obtained evidence. To understand the issue, we trace agentic behaviors into two coupled capabilities: Action Selection (AS), which determines observation streams, and Belief Tracking (BT), which updates the agent's internal task understanding. Theoretical and empirical analyses reveal a bidirectional bottleneck that leads to SeL: weak BT obscures the credit of informative actions, while weak AS deprives BT of useful evidence. This coupling weakens the learning signal for both capabilities and leads to SeL. To mitigate this issue, we propose AREW, a simple yet effective Advantage Reweighting method that uses easy-to-obtain directional critiques to reallocate credit within trajectories. Extensive experiments across 9 agentic tasks of varying complexity show that AREW significantly mitigates SeL, yielding up to 60-point gains in final performance. Code is available at https://github.com/unimpor/T3.
title On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents
topic Artificial Intelligence
url https://arxiv.org/abs/2603.12109