The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jierun, Yu, Tiezheng, Bai, Haoli, Yao, Lewei, Wu, Jiannan, Li, Kaican, Mi, Fei, Tao, Chaofan, Zhu, Lei, Zhang, Manyi, Li, Xiaohui, Hou, Lu, Shang, Lifeng, Liu, Qun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909682585567232
author Chen, Jierun
Yu, Tiezheng
Bai, Haoli
Yao, Lewei
Wu, Jiannan
Li, Kaican
Mi, Fei
Tao, Chaofan
Zhu, Lei
Zhang, Manyi
Li, Xiaohui
Hou, Lu
Shang, Lifeng
Liu, Qun
author_facet Chen, Jierun
Yu, Tiezheng
Bai, Haoli
Yao, Lewei
Wu, Jiannan
Li, Kaican
Mi, Fei
Tao, Chaofan
Zhu, Lei
Zhang, Manyi
Li, Xiaohui
Hou, Lu
Shang, Lifeng
Liu, Qun
contents Large vision-language models (VLMs) increasingly adopt post-training techniques such as long chain-of-thought (CoT) supervised fine-tuning (SFT) and reinforcement learning (RL) to elicit sophisticated reasoning. While these methods exhibit synergy in language-only models, their joint effectiveness in VLMs remains uncertain. We present a systematic investigation into the distinct roles and interplay of long-CoT SFT and RL across multiple multimodal reasoning benchmarks. We find that SFT improves performance on difficult questions by in-depth, structured reasoning, but introduces verbosity and degrades performance on simpler ones. In contrast, RL promotes generalization and brevity, yielding consistent improvements across all difficulty levels, though the improvements on the hardest questions are less prominent compared to SFT. Surprisingly, combining them through two-staged, interleaved, or progressive training strategies, as well as data mixing and model merging, all fails to produce additive benefits, instead leading to trade-offs in accuracy, reasoning style, and response length. This ``synergy dilemma'' highlights the need for more seamless and adaptive approaches to unlock the full potential of combined post-training techniques for reasoning VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07562
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
Chen, Jierun
Yu, Tiezheng
Bai, Haoli
Yao, Lewei
Wu, Jiannan
Li, Kaican
Mi, Fei
Tao, Chaofan
Zhu, Lei
Zhang, Manyi
Li, Xiaohui
Hou, Lu
Shang, Lifeng
Liu, Qun
Computation and Language
Large vision-language models (VLMs) increasingly adopt post-training techniques such as long chain-of-thought (CoT) supervised fine-tuning (SFT) and reinforcement learning (RL) to elicit sophisticated reasoning. While these methods exhibit synergy in language-only models, their joint effectiveness in VLMs remains uncertain. We present a systematic investigation into the distinct roles and interplay of long-CoT SFT and RL across multiple multimodal reasoning benchmarks. We find that SFT improves performance on difficult questions by in-depth, structured reasoning, but introduces verbosity and degrades performance on simpler ones. In contrast, RL promotes generalization and brevity, yielding consistent improvements across all difficulty levels, though the improvements on the hardest questions are less prominent compared to SFT. Surprisingly, combining them through two-staged, interleaved, or progressive training strategies, as well as data mixing and model merging, all fails to produce additive benefits, instead leading to trade-offs in accuracy, reasoning style, and response length. This ``synergy dilemma'' highlights the need for more seamless and adaptive approaches to unlock the full potential of combined post-training techniques for reasoning VLMs.
title The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
topic Computation and Language
url https://arxiv.org/abs/2507.07562