RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914164912422912 |
|---|---|
| author | Xu, Meilong Fu, Di Zhang, Jiaxing Yu, Gong Zheng, Jiayu Hu, Xiaoling Zhao, Dongdi Li, Feiyang Chen, Chao Cao, Yong |
| author_facet | Xu, Meilong Fu, Di Zhang, Jiaxing Yu, Gong Zheng, Jiayu Hu, Xiaoling Zhao, Dongdi Li, Feiyang Chen, Chao Cao, Yong |
| contents | Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particularly in cases with limited data. This stems from a critical \textit{rationale gap}, where sparse domain data is insufficient to bridge the semantic distance between complex spatio-temporal content and abstract classification labels. We propose a two-stage self-improvement paradigm to bridge this gap without new annotations. First, we prompt the VLMs to generate detailed textual rationales for each video, compelling them to articulate the domain-specific logic. The VLM is then fine-tuned on these self-generated rationales, utilizing this intermediate supervision to align its representations with the nuances of the target domain. Second, conventional supervised fine-tuning (SFT) is performed on the task labels, achieving markedly higher effectiveness as a result of the model's pre-acquired domain reasoning. Extensive experiments on diverse datasets demonstrate that our method significantly outperforms direct SFT, validating self-generated rationale as an effective, annotation-efficient paradigm for adapting VLMs to domain-specific video analysis. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_15923 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification Xu, Meilong Fu, Di Zhang, Jiaxing Yu, Gong Zheng, Jiayu Hu, Xiaoling Zhao, Dongdi Li, Feiyang Chen, Chao Cao, Yong Computer Vision and Pattern Recognition Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particularly in cases with limited data. This stems from a critical \textit{rationale gap}, where sparse domain data is insufficient to bridge the semantic distance between complex spatio-temporal content and abstract classification labels. We propose a two-stage self-improvement paradigm to bridge this gap without new annotations. First, we prompt the VLMs to generate detailed textual rationales for each video, compelling them to articulate the domain-specific logic. The VLM is then fine-tuned on these self-generated rationales, utilizing this intermediate supervision to align its representations with the nuances of the target domain. Second, conventional supervised fine-tuning (SFT) is performed on the task labels, achieving markedly higher effectiveness as a result of the model's pre-acquired domain reasoning. Extensive experiments on diverse datasets demonstrate that our method significantly outperforms direct SFT, validating self-generated rationale as an effective, annotation-efficient paradigm for adapting VLMs to domain-specific video analysis. |
| title | RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.15923 |