ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yuqi, Chen, Liangyu, Liu, Jiazhen, Zhu, Mingkang, Zhong, Zhisheng, Yu, Bei, Jia, Jiaya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909030384926720
author Liu, Yuqi
Chen, Liangyu
Liu, Jiazhen
Zhu, Mingkang
Zhong, Zhisheng
Yu, Bei
Jia, Jiaya
author_facet Liu, Yuqi
Chen, Liangyu
Liu, Jiazhen
Zhu, Mingkang
Zhong, Zhisheng
Yu, Bei
Jia, Jiaya
contents Post-training Large Vision-and-Language Models (LVLMs) typically involves Supervised Fine-Tuning (SFT) for knowledge injection or Reinforcement Learning with Verifiable Rewards (RLVR) for performance enhancement. However, SFT often leads to sub-optimal performance, while RLVR remains constrained by the model's internal knowledge base. While a sequential SFT $\rightarrow$ RLVR pipeline can be used, it introduces significant computational overhead and suffers from catastrophic forgetting. To address these limitations, we propose ViSurf (\textbf{Vi}sual \textbf{Su}pervised-and-\textbf{R}einforcement \textbf{F}ine-Tuning), a unified, single-stage paradigm that integrates the strengths of both SFT and RLVR. By analyzing their training objectives, we establish a unified framework that injects ground-truth labels directly into RLVR rollouts, facilitating simultaneous external supervision and internal reinforcement. Furthermore, we introduce three novel reward control strategies to ensure training stability and optimization. Extensive experiments demonstrate that ViSurf consistently outperforms standalone SFT, RLVR, and the traditional two-stage pipeline across diverse benchmarks. In-depth analysis corroborates these findings, validating the derivation and design principles of ViSurf.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10606
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
Liu, Yuqi
Chen, Liangyu
Liu, Jiazhen
Zhu, Mingkang
Zhong, Zhisheng
Yu, Bei
Jia, Jiaya
Computer Vision and Pattern Recognition
Post-training Large Vision-and-Language Models (LVLMs) typically involves Supervised Fine-Tuning (SFT) for knowledge injection or Reinforcement Learning with Verifiable Rewards (RLVR) for performance enhancement. However, SFT often leads to sub-optimal performance, while RLVR remains constrained by the model's internal knowledge base. While a sequential SFT $\rightarrow$ RLVR pipeline can be used, it introduces significant computational overhead and suffers from catastrophic forgetting. To address these limitations, we propose ViSurf (\textbf{Vi}sual \textbf{Su}pervised-and-\textbf{R}einforcement \textbf{F}ine-Tuning), a unified, single-stage paradigm that integrates the strengths of both SFT and RLVR. By analyzing their training objectives, we establish a unified framework that injects ground-truth labels directly into RLVR rollouts, facilitating simultaneous external supervision and internal reinforcement. Furthermore, we introduce three novel reward control strategies to ensure training stability and optimization. Extensive experiments demonstrate that ViSurf consistently outperforms standalone SFT, RLVR, and the traditional two-stage pipeline across diverse benchmarks. In-depth analysis corroborates these findings, validating the derivation and design principles of ViSurf.
title ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.10606