Test-Time Deep Thinking to Explore Implicit Rules

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Wentong, Cong, Xin, Zhang, Zhong, Lu, Yaxi, Zhao, Siyuan, Wu, Yesai, Luo, Qinyu, Chen, Haotian, Lin, Yankai, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917552127475712
author Chen, Wentong
Cong, Xin
Zhang, Zhong
Lu, Yaxi
Zhao, Siyuan
Wu, Yesai
Luo, Qinyu
Chen, Haotian
Lin, Yankai
Liu, Zhiyuan
Sun, Maosong
author_facet Chen, Wentong
Cong, Xin
Zhang, Zhong
Lu, Yaxi
Zhao, Siyuan
Wu, Yesai
Luo, Qinyu
Chen, Haotian
Lin, Yankai
Liu, Zhiyuan
Sun, Maosong
contents With the continuous advancement of Large Language Models (LLMs), intelligent agents are becoming increasingly vital. However, these agents often fail in environments governed by implicit rules--hidden constraints that cannot be observed directly and must be inferred through interaction. This causes agents to fall into repetitive trial-and-error loops, ultimately leading to task failure. To address this challenge, we propose Test-Time Exploration (TTExplore), a framework where a thinker component analyzes interaction history to infer these implicit rules and guide an actor. Effective exploration in this setting critically depends on the reasoning ability of the thinker. However, evaluating deep reasoning trajectories is inherently unstable and difficult, which poses a major obstacle to effective training. To overcome this issue, we introduce a novel and stable reinforcement learning pipeline. The core idea is to use accurate task-level scores as indirect rewards to bypass the difficulty of evaluating intermediate reasoning, and to retain only a single thinking node per trajectory to alleviate reward sparsity. Using this pipeline, we train a specialized 7B model, Exp-Thinker. Experiments on five text-based embodied tasks show that TTExplore equipped with Exp-Thinker improves baseline agent performance by an average of $14$-$19$ points, demonstrating the effectiveness of explicitly reasoning about implicit rules.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24828
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Test-Time Deep Thinking to Explore Implicit Rules
Chen, Wentong
Cong, Xin
Zhang, Zhong
Lu, Yaxi
Zhao, Siyuan
Wu, Yesai
Luo, Qinyu
Chen, Haotian
Lin, Yankai
Liu, Zhiyuan
Sun, Maosong
Artificial Intelligence
With the continuous advancement of Large Language Models (LLMs), intelligent agents are becoming increasingly vital. However, these agents often fail in environments governed by implicit rules--hidden constraints that cannot be observed directly and must be inferred through interaction. This causes agents to fall into repetitive trial-and-error loops, ultimately leading to task failure. To address this challenge, we propose Test-Time Exploration (TTExplore), a framework where a thinker component analyzes interaction history to infer these implicit rules and guide an actor. Effective exploration in this setting critically depends on the reasoning ability of the thinker. However, evaluating deep reasoning trajectories is inherently unstable and difficult, which poses a major obstacle to effective training. To overcome this issue, we introduce a novel and stable reinforcement learning pipeline. The core idea is to use accurate task-level scores as indirect rewards to bypass the difficulty of evaluating intermediate reasoning, and to retain only a single thinking node per trajectory to alleviate reward sparsity. Using this pipeline, we train a specialized 7B model, Exp-Thinker. Experiments on five text-based embodied tasks show that TTExplore equipped with Exp-Thinker improves baseline agent performance by an average of $14$-$19$ points, demonstrating the effectiveness of explicitly reasoning about implicit rules.
title Test-Time Deep Thinking to Explore Implicit Rules
topic Artificial Intelligence
url https://arxiv.org/abs/2605.24828