NFT: Bridging Supervised Learning and Reinforcement Learning in Math Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Huayu, Zheng, Kaiwen, Zhang, Qinsheng, Cui, Ganqu, Yuan, Lifan, Cui, Yin, Ye, Haotian, Lin, Tsung-Yi, Liu, Ming-Yu, Zhu, Jun, Wang, Haoxiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908857495715840
author Chen, Huayu
Zheng, Kaiwen
Zhang, Qinsheng
Cui, Ganqu
Yuan, Lifan
Cui, Yin
Ye, Haotian
Lin, Tsung-Yi
Liu, Ming-Yu
Zhu, Jun
Wang, Haoxiang
author_facet Chen, Huayu
Zheng, Kaiwen
Zhang, Qinsheng
Cui, Ganqu
Yuan, Lifan
Cui, Yin
Ye, Haotian
Lin, Tsung-Yi
Liu, Ming-Yu
Zhu, Jun
Wang, Haoxiang
contents Reinforcement Learning (RL) has played a central role in the recent surge of LLMs' math abilities by enabling self-improvement through binary verifier signals. In contrast, Supervised Learning (SL) is rarely considered for such verification-driven training, largely due to its heavy reliance on reference answers and inability to reflect on mistakes. In this work, we challenge the prevailing notion that self-improvement is exclusive to RL and propose Negative-aware Fine-Tuning (NFT) -- a supervised approach that enables LLMs to reflect on their failures and improve autonomously with no external teachers. In online training, instead of throwing away self-generated negative answers, NFT constructs an implicit negative policy to model them. This implicit policy is parameterized with the same positive LLM we target to optimize on positive data, enabling direct policy optimization on all LLMs' generations. We conduct experiments on 7B and 32B models in math reasoning tasks. Results consistently show that through the additional leverage of negative feedback, NFT significantly improves over SL baselines like Rejection sampling Fine-Tuning, matching or even surpassing leading RL algorithms like GRPO and DAPO. Furthermore, we demonstrate that NFT and GRPO are actually equivalent in strict-on-policy training, even though they originate from entirely different theoretical foundations. Our experiments and theoretical findings bridge the gap between SL and RL methods in binary-feedback learning systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18116
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NFT: Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
Chen, Huayu
Zheng, Kaiwen
Zhang, Qinsheng
Cui, Ganqu
Yuan, Lifan
Cui, Yin
Ye, Haotian
Lin, Tsung-Yi
Liu, Ming-Yu
Zhu, Jun
Wang, Haoxiang
Machine Learning
Computation and Language
Reinforcement Learning (RL) has played a central role in the recent surge of LLMs' math abilities by enabling self-improvement through binary verifier signals. In contrast, Supervised Learning (SL) is rarely considered for such verification-driven training, largely due to its heavy reliance on reference answers and inability to reflect on mistakes. In this work, we challenge the prevailing notion that self-improvement is exclusive to RL and propose Negative-aware Fine-Tuning (NFT) -- a supervised approach that enables LLMs to reflect on their failures and improve autonomously with no external teachers. In online training, instead of throwing away self-generated negative answers, NFT constructs an implicit negative policy to model them. This implicit policy is parameterized with the same positive LLM we target to optimize on positive data, enabling direct policy optimization on all LLMs' generations. We conduct experiments on 7B and 32B models in math reasoning tasks. Results consistently show that through the additional leverage of negative feedback, NFT significantly improves over SL baselines like Rejection sampling Fine-Tuning, matching or even surpassing leading RL algorithms like GRPO and DAPO. Furthermore, we demonstrate that NFT and GRPO are actually equivalent in strict-on-policy training, even though they originate from entirely different theoretical foundations. Our experiments and theoretical findings bridge the gap between SL and RL methods in binary-feedback learning systems.
title NFT: Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2505.18116