First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Lai, Li, Yuting, Wang, Chen, Wang, Yue, Kong, Linghe, Huang, Weiran, Sun, Lichao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
by: Li, Yuting, et al.
Published: (2025)
by: Li, Yuting, et al.
Published: (2025)
IDER: IDempotent Experience Replay for Reliable Continual Learning
by: Liu, Zhanwang, et al.
Published: (2026)
by: Liu, Zhanwang, et al.
Published: (2026)
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
by: Wei, Lai, et al.
Published: (2023)
by: Wei, Lai, et al.
Published: (2023)
DeReason: A Difficulty-Aware Curriculum Improves Decoupled SFT-then-RL Training for General Reasoning
by: Hu, Hanxu, et al.
Published: (2026)
by: Hu, Hanxu, et al.
Published: (2026)
The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
by: Chen, Jierun, et al.
Published: (2025)
by: Chen, Jierun, et al.
Published: (2025)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
by: Xu, Chengzhi, et al.
Published: (2025)
by: Xu, Chengzhi, et al.
Published: (2025)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
by: Limozin, Alexis, et al.
Published: (2026)
by: Limozin, Alexis, et al.
Published: (2026)
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
by: Chen, Hardy, et al.
Published: (2025)
by: Chen, Hardy, et al.
Published: (2025)
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
by: Deng, Yihe, et al.
Published: (2025)
by: Deng, Yihe, et al.
Published: (2025)
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
by: Lu, Aojun, et al.
Published: (2026)
by: Lu, Aojun, et al.
Published: (2026)
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Parallel-SFT: Improving Zero-Shot Cross-Programming-Language Transfer for Code RL
by: Wu, Zhaofeng, et al.
Published: (2026)
by: Wu, Zhaofeng, et al.
Published: (2026)
FinLLM-B: When Large Language Models Meet Financial Breakout Trading
by: Zhang, Kang, et al.
Published: (2024)
by: Zhang, Kang, et al.
Published: (2024)
LLM4Brain: Training a Large Language Model for Brain Video Understanding
by: Zheng, Ruizhe, et al.
Published: (2024)
by: Zheng, Ruizhe, et al.
Published: (2024)
Continual SFT Matches Multimodal RLHF with Negative Supervision
by: Zhu, Ke, et al.
Published: (2024)
by: Zhu, Ke, et al.
Published: (2024)
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
by: Wang, Sudong, et al.
Published: (2026)
by: Wang, Sudong, et al.
Published: (2026)
TransMed: Large Language Models Enhance Vision Transformer for Biomedical Image Classification
by: Zheng, Kaipeng, et al.
Published: (2023)
by: Zheng, Kaipeng, et al.
Published: (2023)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
by: Qiu, Haibo, et al.
Published: (2025)
by: Qiu, Haibo, et al.
Published: (2025)
Enhanced Continual Learning of Vision-Language Models with Model Fusion
by: Gao, Haoyuan, et al.
Published: (2025)
by: Gao, Haoyuan, et al.
Published: (2025)
Improving SAM for Camouflaged Object Detection via Dual Stream Adapters
by: Liu, Jiaming, et al.
Published: (2025)
by: Liu, Jiaming, et al.
Published: (2025)
Learning to Adapt SFT Data for Better Reasoning Generalization
by: Sun, Lisong, et al.
Published: (2026)
by: Sun, Lisong, et al.
Published: (2026)
$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training
by: Zhou, Jin Peng, et al.
Published: (2025)
by: Zhou, Jin Peng, et al.
Published: (2025)
Horizon-LM: A RAM-Centric Architecture for LLM Training
by: Yuan, Zhengqing, et al.
Published: (2026)
by: Yuan, Zhengqing, et al.
Published: (2026)
GenCLS++: Pushing the Boundaries of Generative Classification in LLMs Through Comprehensive SFT and RL Studies Across Diverse Datasets
by: He, Mingqian, et al.
Published: (2025)
by: He, Mingqian, et al.
Published: (2025)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
by: Zhang, Charlie, et al.
Published: (2025)
by: Zhang, Charlie, et al.
Published: (2025)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
by: Chu, Tianzhe, et al.
Published: (2025)
by: Chu, Tianzhe, et al.
Published: (2025)
From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection
by: Li, Youpeng, et al.
Published: (2026)
by: Li, Youpeng, et al.
Published: (2026)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
by: Wei, Lai, et al.
Published: (2026)
by: Wei, Lai, et al.
Published: (2026)
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
by: Liu, Zihan, et al.
Published: (2025)
by: Liu, Zihan, et al.
Published: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
by: Liu, Yixin, et al.
Published: (2026)
by: Liu, Yixin, et al.
Published: (2026)
SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
by: Wang, Junke, et al.
Published: (2025)
by: Wang, Junke, et al.
Published: (2025)
Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT
by: Wang, Jiacheng, et al.
Published: (2026)
by: Wang, Jiacheng, et al.
Published: (2026)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
On Designing Effective RL Reward at Training Time for LLM Reasoning
by: Gao, Jiaxuan, et al.
Published: (2024)
by: Gao, Jiaxuan, et al.
Published: (2024)
SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization
by: Su, Xiaole, et al.
Published: (2026)
by: Su, Xiaole, et al.
Published: (2026)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
by: Zheng, Haojie, et al.
Published: (2024)
by: Zheng, Haojie, et al.
Published: (2024)
When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains
by: Jeddi, Ahmadreza, et al.
Published: (2026)
by: Jeddi, Ahmadreza, et al.
Published: (2026)
Similar Items
-
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
by: Wei, Lai, et al.
Published: (2025) -
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
by: Li, Yuting, et al.
Published: (2025) -
IDER: IDempotent Experience Replay for Reliable Continual Learning
by: Liu, Zhanwang, et al.
Published: (2026) -
Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning
by: Chen, Liang, et al.
Published: (2025) -
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
by: Wei, Lai, et al.
Published: (2023)