Reasoning with Reinforced Functional Token Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kongcheng, Yao, Qi, Lai, Baisheng, Huang, Jiaxing, Fang, Wenkai, Tao, Dacheng, Song, Mingli, Liu, Shunyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915160650678272
author Zhang, Kongcheng
Yao, Qi
Lai, Baisheng
Huang, Jiaxing
Fang, Wenkai
Tao, Dacheng
Song, Mingli
Liu, Shunyu
author_facet Zhang, Kongcheng
Yao, Qi
Lai, Baisheng
Huang, Jiaxing
Fang, Wenkai
Tao, Dacheng
Song, Mingli
Liu, Shunyu
contents In this work, we propose Reinforced Functional Token Tuning (RFTT), a novel reinforced fine-tuning framework that empowers Large Language Models (LLMs) with self-play learn-to-reason capabilities. Unlike prior prompt-driven reasoning efforts, RFTT embeds a rich set of learnable functional tokens (e.g., <analyze>, <verify>, <refine>) directly into the model vocabulary, enabling chain-of-thought construction with diverse human-like reasoning behaviors. Specifically, RFTT comprises two phases: (1) supervised fine-tuning performs prompt-driven tree search to obtain self-generated training data annotated with functional tokens, which warms up the model to learn these tokens for reasoning; and (2) online reinforcement learning further allows the model to explore different reasoning pathways through functional token sampling without relying on prompts, thereby facilitating effective self-improvement for functional reasoning. Extensive experiments demonstrate the superiority of the proposed RFTT on mathematical benchmarks, significantly boosting Qwen-2.5-7B-Instruct (70.6% to 79.8%) and LLaMA-3.1-8B-Instruct (32.2% to 60.2%) on the MATH dataset. Moreover, the performance of RFTT consistently improves with more search rollouts at inference time. Our code is available at https://github.com/sastpg/RFTT.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13389
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reasoning with Reinforced Functional Token Tuning
Zhang, Kongcheng
Yao, Qi
Lai, Baisheng
Huang, Jiaxing
Fang, Wenkai
Tao, Dacheng
Song, Mingli
Liu, Shunyu
Artificial Intelligence
In this work, we propose Reinforced Functional Token Tuning (RFTT), a novel reinforced fine-tuning framework that empowers Large Language Models (LLMs) with self-play learn-to-reason capabilities. Unlike prior prompt-driven reasoning efforts, RFTT embeds a rich set of learnable functional tokens (e.g., <analyze>, <verify>, <refine>) directly into the model vocabulary, enabling chain-of-thought construction with diverse human-like reasoning behaviors. Specifically, RFTT comprises two phases: (1) supervised fine-tuning performs prompt-driven tree search to obtain self-generated training data annotated with functional tokens, which warms up the model to learn these tokens for reasoning; and (2) online reinforcement learning further allows the model to explore different reasoning pathways through functional token sampling without relying on prompts, thereby facilitating effective self-improvement for functional reasoning. Extensive experiments demonstrate the superiority of the proposed RFTT on mathematical benchmarks, significantly boosting Qwen-2.5-7B-Instruct (70.6% to 79.8%) and LLaMA-3.1-8B-Instruct (32.2% to 60.2%) on the MATH dataset. Moreover, the performance of RFTT consistently improves with more search rollouts at inference time. Our code is available at https://github.com/sastpg/RFTT.
title Reasoning with Reinforced Functional Token Tuning
topic Artificial Intelligence
url https://arxiv.org/abs/2502.13389