Control-R: Towards controllable test-time scaling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Di, Wang, Weida, Li, Junxian, Wang, Xunzhi, Li, Jiatong, Wu, Jianbo, Lei, Jingdi, He, Haonan, Ye, Peng, Zhang, Shufei, Ouyang, Wanli, Li, Yuqiang, Zhou, Dongzhan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908387850059776
author Zhang, Di
Wang, Weida
Li, Junxian
Wang, Xunzhi
Li, Jiatong
Wu, Jianbo
Lei, Jingdi
He, Haonan
Ye, Peng
Zhang, Shufei
Ouyang, Wanli
Li, Yuqiang
Zhou, Dongzhan
author_facet Zhang, Di
Wang, Weida
Li, Junxian
Wang, Xunzhi
Li, Jiatong
Wu, Jianbo
Lei, Jingdi
He, Haonan
Ye, Peng
Zhang, Shufei
Ouyang, Wanli
Li, Yuqiang
Zhou, Dongzhan
contents This paper target in addressing the challenges of underthinking and overthinking in long chain-of-thought (CoT) reasoning for Large Reasoning Models (LRMs) by introducing Reasoning Control Fields (RCF)--a novel test-time approach that injects structured control signals to guide reasoning from a tree search perspective. RCF enables models to adjust reasoning effort according to given control conditions when solving complex tasks. Additionally, we present the Control-R-4K dataset, which consists of challenging problems annotated with detailed reasoning processes and corresponding control fields. To further enhance reasoning control, we propose a Conditional Distillation Finetuning (CDF) method, which trains model--particularly Control-R-32B--to effectively adjust reasoning effort during test time. Experimental results on benchmarks such as AIME2024 and MATH500 demonstrate that our approach achieves state-of-the-art performance at the 32B scale while enabling a controllable Long CoT reasoning process (L-CoT). Overall, this work introduces an effective paradigm for controllable test-time scaling reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00189
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Control-R: Towards controllable test-time scaling
Zhang, Di
Wang, Weida
Li, Junxian
Wang, Xunzhi
Li, Jiatong
Wu, Jianbo
Lei, Jingdi
He, Haonan
Ye, Peng
Zhang, Shufei
Ouyang, Wanli
Li, Yuqiang
Zhou, Dongzhan
Artificial Intelligence
Computation and Language
This paper target in addressing the challenges of underthinking and overthinking in long chain-of-thought (CoT) reasoning for Large Reasoning Models (LRMs) by introducing Reasoning Control Fields (RCF)--a novel test-time approach that injects structured control signals to guide reasoning from a tree search perspective. RCF enables models to adjust reasoning effort according to given control conditions when solving complex tasks. Additionally, we present the Control-R-4K dataset, which consists of challenging problems annotated with detailed reasoning processes and corresponding control fields. To further enhance reasoning control, we propose a Conditional Distillation Finetuning (CDF) method, which trains model--particularly Control-R-32B--to effectively adjust reasoning effort during test time. Experimental results on benchmarks such as AIME2024 and MATH500 demonstrate that our approach achieves state-of-the-art performance at the 32B scale while enabling a controllable Long CoT reasoning process (L-CoT). Overall, this work introduces an effective paradigm for controllable test-time scaling reasoning.
title Control-R: Towards controllable test-time scaling
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.00189