RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zehua, Liu, Shuqi, Zhong, Tao, Yuan, Mingxuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918463742672896
author Liu, Zehua
Liu, Shuqi
Zhong, Tao
Yuan, Mingxuan
author_facet Liu, Zehua
Liu, Shuqi
Zhong, Tao
Yuan, Mingxuan
contents While Supervised Fine-Tuning (SFT) and Rejection Sampling Fine-Tuning (RFT) are standard for LLM alignment, they either rely on costly expert data or discard valuable negative samples, leading to data inefficiency. To address this, we propose Reward Informed Fine-Tuning (RIFT), a simple yet effective framework that utilizes all self-generated samples. Unlike the hard thresholding of RFT, RIFT repurposes negative trajectories, reweighting the loss with scalar rewards to learn from both the positive and negative trajectories from the model outputs. To overcome the training collapse caused by naive reward integration, where direct multiplication yields an unbounded loss, we introduce a stabilized loss formulation that ensures numerical robustness and optimization efficiency. Extensive experiments on mathematical benchmarks across various base models show that RIFT consistently outperforms RFT. Our results demonstrate that RIFT is a robust and data-efficient alternative for alignment using mixed-quality, self-generated data.
format Preprint
id arxiv_https___arxiv_org_abs_2601_09253
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning
Liu, Zehua
Liu, Shuqi
Zhong, Tao
Yuan, Mingxuan
Machine Learning
Artificial Intelligence
While Supervised Fine-Tuning (SFT) and Rejection Sampling Fine-Tuning (RFT) are standard for LLM alignment, they either rely on costly expert data or discard valuable negative samples, leading to data inefficiency. To address this, we propose Reward Informed Fine-Tuning (RIFT), a simple yet effective framework that utilizes all self-generated samples. Unlike the hard thresholding of RFT, RIFT repurposes negative trajectories, reweighting the loss with scalar rewards to learn from both the positive and negative trajectories from the model outputs. To overcome the training collapse caused by naive reward integration, where direct multiplication yields an unbounded loss, we introduce a stabilized loss formulation that ensures numerical robustness and optimization efficiency. Extensive experiments on mathematical benchmarks across various base models show that RIFT consistently outperforms RFT. Our results demonstrate that RIFT is a robust and data-efficient alternative for alignment using mixed-quality, self-generated data.
title RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2601.09253