VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meng, Desen, Huang, Rui, Dai, Zhilin, Li, Xinhao, Xu, Yifan, Zhang, Jun, Huang, Zhenpeng, Zhang, Meng, Zhang, Lingshu, Liu, Yi, Wang, Limin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915318315614208
author Meng, Desen
Huang, Rui
Dai, Zhilin
Li, Xinhao
Xu, Yifan
Zhang, Jun
Huang, Zhenpeng
Zhang, Meng
Zhang, Lingshu
Liu, Yi
Wang, Limin
author_facet Meng, Desen
Huang, Rui
Dai, Zhilin
Li, Xinhao
Xu, Yifan
Zhang, Jun
Huang, Zhenpeng
Zhang, Meng
Zhang, Lingshu
Liu, Yi
Wang, Limin
contents While recent advances in reinforcement learning have significantly enhanced reasoning capabilities in large language models (LLMs), these techniques remain underexplored in multi-modal LLMs for video captioning. This paper presents the first systematic investigation of GRPO-based RL post-training for video MLLMs, with the goal of enhancing video MLLMs' capability of describing actions in videos. Specifically, we develop the VideoCap-R1, which is prompted to first perform structured thinking that analyzes video subjects with their attributes and actions before generating complete captions, supported by two specialized reward mechanisms: a LLM-free think scorer evaluating the structured thinking quality and a LLM-assisted caption scorer assessing the output quality. The RL training framework effectively establishes the connection between structured reasoning and comprehensive description generation, enabling the model to produce captions with more accurate actions. Our experiments demonstrate that VideoCap-R1 achieves substantial improvements over the Qwen2VL-7B baseline using limited samples (1.5k) across multiple video caption benchmarks (DREAM1K: +4.4 event F1, VDC: +4.2 Acc, CAREBENCH: +3.1 action F1, +6.9 object F1) while consistently outperforming the SFT-trained counterparts, confirming GRPO's superiority in enhancing MLLMs' captioning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01725
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Meng, Desen
Huang, Rui
Dai, Zhilin
Li, Xinhao
Xu, Yifan
Zhang, Jun
Huang, Zhenpeng
Zhang, Meng
Zhang, Lingshu
Liu, Yi
Wang, Limin
Computer Vision and Pattern Recognition
While recent advances in reinforcement learning have significantly enhanced reasoning capabilities in large language models (LLMs), these techniques remain underexplored in multi-modal LLMs for video captioning. This paper presents the first systematic investigation of GRPO-based RL post-training for video MLLMs, with the goal of enhancing video MLLMs' capability of describing actions in videos. Specifically, we develop the VideoCap-R1, which is prompted to first perform structured thinking that analyzes video subjects with their attributes and actions before generating complete captions, supported by two specialized reward mechanisms: a LLM-free think scorer evaluating the structured thinking quality and a LLM-assisted caption scorer assessing the output quality. The RL training framework effectively establishes the connection between structured reasoning and comprehensive description generation, enabling the model to produce captions with more accurate actions. Our experiments demonstrate that VideoCap-R1 achieves substantial improvements over the Qwen2VL-7B baseline using limited samples (1.5k) across multiple video caption benchmarks (DREAM1K: +4.4 event F1, VDC: +4.2 Acc, CAREBENCH: +3.1 action F1, +6.9 object F1) while consistently outperforming the SFT-trained counterparts, confirming GRPO's superiority in enhancing MLLMs' captioning capabilities.
title VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.01725