Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ren, Qihan, Wang, Peng, Cai, Ruikun, Shao, Shuai, Guo, Dadi, Xie, Yuejin, Li, Yafu, Zhang, Quanshi, Hu, Xia, Shao, Jing, Liu, Dongrui
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918433538441216
author Ren, Qihan
Wang, Peng
Cai, Ruikun
Shao, Shuai
Guo, Dadi
Xie, Yuejin
Li, Yafu
Zhang, Quanshi
Hu, Xia
Shao, Jing
Liu, Dongrui
author_facet Ren, Qihan
Wang, Peng
Cai, Ruikun
Shao, Shuai
Guo, Dadi
Xie, Yuejin
Li, Yafu
Zhang, Quanshi
Hu, Xia
Shao, Jing
Liu, Dongrui
contents A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit this claim for reasoning SFT with long chain-of-thought (CoT) supervision and find that cross-domain generalization is not absent but conditional, jointly shaped by optimization dynamics, training data, and base-model capability. Some reported failures are under-optimization artifacts: cross-domain performance first degrades before recovering and improving with extended training (a dip-and-recovery pattern), so shorttraining checkpoints can underestimate generalization. Data quality and structure both matter: low-quality solutions broadly hurt generalization,while verified long-CoT traces yield consistent cross-domain gains. Model capability is essential: stronger models internalize transferable procedural patterns (e.g., backtracking) even from a toy arithmetic game, while weaker ones imitate surface verbosity. This generalization is asymmetric, however: reasoning improves while safety degrades, reframing the question from whether reasoning SFT generalizes to under what conditions and at what cost.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06628
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
Ren, Qihan
Wang, Peng
Cai, Ruikun
Shao, Shuai
Guo, Dadi
Xie, Yuejin
Li, Yafu
Zhang, Quanshi
Hu, Xia
Shao, Jing
Liu, Dongrui
Artificial Intelligence
A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit this claim for reasoning SFT with long chain-of-thought (CoT) supervision and find that cross-domain generalization is not absent but conditional, jointly shaped by optimization dynamics, training data, and base-model capability. Some reported failures are under-optimization artifacts: cross-domain performance first degrades before recovering and improving with extended training (a dip-and-recovery pattern), so shorttraining checkpoints can underestimate generalization. Data quality and structure both matter: low-quality solutions broadly hurt generalization,while verified long-CoT traces yield consistent cross-domain gains. Model capability is essential: stronger models internalize transferable procedural patterns (e.g., backtracking) even from a toy arithmetic game, while weaker ones imitate surface verbosity. This generalization is asymmetric, however: reasoning improves while safety degrades, reframing the question from whether reasoning SFT generalizes to under what conditions and at what cost.
title Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
topic Artificial Intelligence
url https://arxiv.org/abs/2604.06628