Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Shuyao, Peng, Cheng, Long, Jiangxuan, Xu, Weidi, Chu, Wei, Qi, Yuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917145837830144
author Xu, Shuyao
Peng, Cheng
Long, Jiangxuan
Xu, Weidi
Chu, Wei
Qi, Yuan
author_facet Xu, Shuyao
Peng, Cheng
Long, Jiangxuan
Xu, Weidi
Chu, Wei
Qi, Yuan
contents Recent advances in model distillation show that data from advanced reasoning models can effectively train smaller student models. However, standard practices discard incorrect reasoning traces -- valuable, yet underutilized data. This paper addresses the critical question: How can both positive and negative distilled reasoning traces be effectively leveraged to maximize LLM reasoning performance in an offline setting? We employ a two-stage training recipe: first, Supervised Fine-Tuning (SFT) on positive traces, followed by a refinement stage using both positive and negative traces. We find that a simple REINFORCE-style objective, which we term the Reinforcement Distillation (REDI) objective, outperforms established preference optimization methods like DPO and SimPO in this distillation context. Our empirical evaluations demonstrate the effectiveness of this approach. Notably, our Qwen-REDI-1.5B model, trained on just 131k traces from the open Open-R1 dataset, achieves an 83.1% score on MATH-500. Its performance matches that of DeepSeek-R1-Distill-Qwen-1.5B, a model trained on 800k proprietary data. This result showcases the remarkable data efficiency of utilizing previously discarded negative traces.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24850
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning
Xu, Shuyao
Peng, Cheng
Long, Jiangxuan
Xu, Weidi
Chu, Wei
Qi, Yuan
Machine Learning
Artificial Intelligence
Computation and Language
I.2.6; I.2.7
Recent advances in model distillation show that data from advanced reasoning models can effectively train smaller student models. However, standard practices discard incorrect reasoning traces -- valuable, yet underutilized data. This paper addresses the critical question: How can both positive and negative distilled reasoning traces be effectively leveraged to maximize LLM reasoning performance in an offline setting? We employ a two-stage training recipe: first, Supervised Fine-Tuning (SFT) on positive traces, followed by a refinement stage using both positive and negative traces. We find that a simple REINFORCE-style objective, which we term the Reinforcement Distillation (REDI) objective, outperforms established preference optimization methods like DPO and SimPO in this distillation context. Our empirical evaluations demonstrate the effectiveness of this approach. Notably, our Qwen-REDI-1.5B model, trained on just 131k traces from the open Open-R1 dataset, achieves an 83.1% score on MATH-500. Its performance matches that of DeepSeek-R1-Distill-Qwen-1.5B, a model trained on 800k proprietary data. This result showcases the remarkable data efficiency of utilizing previously discarded negative traces.
title Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning
topic Machine Learning
Artificial Intelligence
Computation and Language
I.2.6; I.2.7
url https://arxiv.org/abs/2505.24850