Mind the Gap: Data Rewriting for Stable Off-Policy Supervised Fine-Tuning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Shiwan, Zhao, Xuyang, Zhou, Jiaming, Kong, Aobo, Li, Qicheng, Qin, Yong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914046833328128
author Zhao, Shiwan
Zhao, Xuyang
Zhou, Jiaming
Kong, Aobo
Li, Qicheng
Qin, Yong
author_facet Zhao, Shiwan
Zhao, Xuyang
Zhou, Jiaming
Kong, Aobo
Li, Qicheng
Qin, Yong
contents Supervised fine-tuning (SFT) of large language models can be viewed as an off-policy learning problem, where expert demonstrations come from a fixed behavior policy while training aims to optimize a target policy. Importance sampling is the standard tool for correcting this distribution mismatch, but large policy gaps lead to skewed weights, high variance, and unstable optimization. Existing methods mitigate this issue with KL penalties or clipping, which passively restrict updates rather than actively reducing the gap. We propose a simple yet effective data rewriting framework that proactively shrinks the policy gap before training. For each problem, correct model-generated solutions are kept as on-policy data, while incorrect ones are rewritten through guided re-solving, falling back to expert demonstrations only when needed. This aligns the training distribution with the target policy, reducing variance and improving stability. To handle residual mismatch after rewriting, we additionally apply importance sampling during training, forming a two-stage approach that combines data-level alignment with lightweight optimization-level correction. Experiments on five mathematical reasoning benchmarks show consistent and significant gains over both vanilla SFT and the state-of-the-art Dynamic Fine-Tuning (DFT) approach. Data and code will be released at https://github.com/NKU-HLT/Off-Policy-SFT.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15157
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mind the Gap: Data Rewriting for Stable Off-Policy Supervised Fine-Tuning
Zhao, Shiwan
Zhao, Xuyang
Zhou, Jiaming
Kong, Aobo
Li, Qicheng
Qin, Yong
Machine Learning
Computation and Language
Supervised fine-tuning (SFT) of large language models can be viewed as an off-policy learning problem, where expert demonstrations come from a fixed behavior policy while training aims to optimize a target policy. Importance sampling is the standard tool for correcting this distribution mismatch, but large policy gaps lead to skewed weights, high variance, and unstable optimization. Existing methods mitigate this issue with KL penalties or clipping, which passively restrict updates rather than actively reducing the gap. We propose a simple yet effective data rewriting framework that proactively shrinks the policy gap before training. For each problem, correct model-generated solutions are kept as on-policy data, while incorrect ones are rewritten through guided re-solving, falling back to expert demonstrations only when needed. This aligns the training distribution with the target policy, reducing variance and improving stability. To handle residual mismatch after rewriting, we additionally apply importance sampling during training, forming a two-stage approach that combines data-level alignment with lightweight optimization-level correction. Experiments on five mathematical reasoning benchmarks show consistent and significant gains over both vanilla SFT and the state-of-the-art Dynamic Fine-Tuning (DFT) approach. Data and code will be released at https://github.com/NKU-HLT/Off-Policy-SFT.
title Mind the Gap: Data Rewriting for Stable Off-Policy Supervised Fine-Tuning
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2509.15157