Test-time Recursive Thinking: Self-Improvement without External Feedback

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhuang, Yufan, Singh, Chandan, Liu, Liyuan, Shen, Yelong, Zhang, Dinghuai, Shang, Jingbo, Gao, Jianfeng, Chen, Weizhu
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914303793168384
author Zhuang, Yufan
Singh, Chandan
Liu, Liyuan
Shen, Yelong
Zhang, Dinghuai
Shang, Jingbo
Gao, Jianfeng
Chen, Weizhu
author_facet Zhuang, Yufan
Singh, Chandan
Liu, Liyuan
Shen, Yelong
Zhang, Dinghuai
Shang, Jingbo
Gao, Jianfeng
Chen, Weizhu
contents Modern Large Language Models (LLMs) have shown rapid improvements in reasoning capabilities, driven largely by reinforcement learning (RL) with verifiable rewards. Here, we ask whether these LLMs can self-improve without the need for additional training. We identify two core challenges for such systems: (i) efficiently generating diverse, high-quality candidate solutions, and (ii) reliably selecting correct answers in the absence of ground-truth supervision. To address these challenges, we propose Test-time Recursive Thinking (TRT), an iterative self-improvement framework that conditions generation on rollout-specific strategies, accumulated knowledge, and self-generated verification signals. Using TRT, open-source models reach 100% accuracy on AIME-25/24, and on LiveCodeBench's most difficult problems, closed-source models improve by 10.4-14.8 percentage points without external feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03094
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Test-time Recursive Thinking: Self-Improvement without External Feedback
Zhuang, Yufan
Singh, Chandan
Liu, Liyuan
Shen, Yelong
Zhang, Dinghuai
Shang, Jingbo
Gao, Jianfeng
Chen, Weizhu
Computation and Language
Modern Large Language Models (LLMs) have shown rapid improvements in reasoning capabilities, driven largely by reinforcement learning (RL) with verifiable rewards. Here, we ask whether these LLMs can self-improve without the need for additional training. We identify two core challenges for such systems: (i) efficiently generating diverse, high-quality candidate solutions, and (ii) reliably selecting correct answers in the absence of ground-truth supervision. To address these challenges, we propose Test-time Recursive Thinking (TRT), an iterative self-improvement framework that conditions generation on rollout-specific strategies, accumulated knowledge, and self-generated verification signals. Using TRT, open-source models reach 100% accuracy on AIME-25/24, and on LiveCodeBench's most difficult problems, closed-source models improve by 10.4-14.8 percentage points without external feedback.
title Test-time Recursive Thinking: Self-Improvement without External Feedback
topic Computation and Language
url https://arxiv.org/abs/2602.03094