RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Meilong, Fu, Di, Zhang, Jiaxing, Yu, Gong, Zheng, Jiayu, Hu, Xiaoling, Zhao, Dongdi, Li, Feiyang, Chen, Chao, Cao, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914164912422912
author Xu, Meilong
Fu, Di
Zhang, Jiaxing
Yu, Gong
Zheng, Jiayu
Hu, Xiaoling
Zhao, Dongdi
Li, Feiyang
Chen, Chao
Cao, Yong
author_facet Xu, Meilong
Fu, Di
Zhang, Jiaxing
Yu, Gong
Zheng, Jiayu
Hu, Xiaoling
Zhao, Dongdi
Li, Feiyang
Chen, Chao
Cao, Yong
contents Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particularly in cases with limited data. This stems from a critical \textit{rationale gap}, where sparse domain data is insufficient to bridge the semantic distance between complex spatio-temporal content and abstract classification labels. We propose a two-stage self-improvement paradigm to bridge this gap without new annotations. First, we prompt the VLMs to generate detailed textual rationales for each video, compelling them to articulate the domain-specific logic. The VLM is then fine-tuned on these self-generated rationales, utilizing this intermediate supervision to align its representations with the nuances of the target domain. Second, conventional supervised fine-tuning (SFT) is performed on the task labels, achieving markedly higher effectiveness as a result of the model's pre-acquired domain reasoning. Extensive experiments on diverse datasets demonstrate that our method significantly outperforms direct SFT, validating self-generated rationale as an effective, annotation-efficient paradigm for adapting VLMs to domain-specific video analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15923
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification
Xu, Meilong
Fu, Di
Zhang, Jiaxing
Yu, Gong
Zheng, Jiayu
Hu, Xiaoling
Zhao, Dongdi
Li, Feiyang
Chen, Chao
Cao, Yong
Computer Vision and Pattern Recognition
Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particularly in cases with limited data. This stems from a critical \textit{rationale gap}, where sparse domain data is insufficient to bridge the semantic distance between complex spatio-temporal content and abstract classification labels. We propose a two-stage self-improvement paradigm to bridge this gap without new annotations. First, we prompt the VLMs to generate detailed textual rationales for each video, compelling them to articulate the domain-specific logic. The VLM is then fine-tuned on these self-generated rationales, utilizing this intermediate supervision to align its representations with the nuances of the target domain. Second, conventional supervised fine-tuning (SFT) is performed on the task labels, achieving markedly higher effectiveness as a result of the model's pre-acquired domain reasoning. Extensive experiments on diverse datasets demonstrate that our method significantly outperforms direct SFT, validating self-generated rationale as an effective, annotation-efficient paradigm for adapting VLMs to domain-specific video analysis.
title RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.15923