Video-R1: Reinforcing Video Reasoning in MLLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Kaituo, Gong, Kaixiong, Li, Bohao, Guo, Zonghao, Wang, Yibing, Peng, Tianshuo, Wu, Junfei, Zhang, Xiaoying, Wang, Benyou, Yue, Xiangyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915569876336640
author Feng, Kaituo
Gong, Kaixiong
Li, Bohao
Guo, Zonghao
Wang, Yibing
Peng, Tianshuo
Wu, Junfei
Zhang, Xiaoying
Wang, Benyou
Yue, Xiangyu
author_facet Feng, Kaituo
Gong, Kaixiong
Li, Bohao
Guo, Zonghao
Wang, Yibing
Peng, Tianshuo
Wu, Junfei
Zhang, Xiaoying
Wang, Benyou
Yue, Xiangyu
contents Inspired by DeepSeek-R1's success in eliciting reasoning abilities through rule-based reinforcement learning (RL), we introduce Video-R1 as the first attempt to systematically explore the R1 paradigm for incentivizing video reasoning within multimodal large language models (MLLMs). However, directly applying RL training with the GRPO algorithm to video reasoning presents two primary challenges: (i) a lack of temporal modeling for video reasoning, and (ii) the scarcity of high-quality video-reasoning data. To address these issues, we first propose the T-GRPO algorithm, which encourages models to utilize temporal information in videos for reasoning. Additionally, instead of relying solely on video data, we incorporate high-quality image-reasoning data into the training process. We have constructed two datasets: Video-R1-CoT-165k for SFT cold start and Video-R1-260k for RL training, both comprising image and video data. Experimental results demonstrate that Video-R1 achieves significant improvements on video reasoning benchmarks such as VideoMMMU and VSI-Bench, as well as on general video benchmarks including MVBench and TempCompass, etc. Notably, Video-R1-7B attains a 37.1% accuracy on video spatial reasoning benchmark VSI-bench, surpassing the commercial proprietary model GPT-4o. All code, models, and data are released in: https://github.com/tulerfeng/Video-R1.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21776
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Video-R1: Reinforcing Video Reasoning in MLLMs
Feng, Kaituo
Gong, Kaixiong
Li, Bohao
Guo, Zonghao
Wang, Yibing
Peng, Tianshuo
Wu, Junfei
Zhang, Xiaoying
Wang, Benyou
Yue, Xiangyu
Computer Vision and Pattern Recognition
Inspired by DeepSeek-R1's success in eliciting reasoning abilities through rule-based reinforcement learning (RL), we introduce Video-R1 as the first attempt to systematically explore the R1 paradigm for incentivizing video reasoning within multimodal large language models (MLLMs). However, directly applying RL training with the GRPO algorithm to video reasoning presents two primary challenges: (i) a lack of temporal modeling for video reasoning, and (ii) the scarcity of high-quality video-reasoning data. To address these issues, we first propose the T-GRPO algorithm, which encourages models to utilize temporal information in videos for reasoning. Additionally, instead of relying solely on video data, we incorporate high-quality image-reasoning data into the training process. We have constructed two datasets: Video-R1-CoT-165k for SFT cold start and Video-R1-260k for RL training, both comprising image and video data. Experimental results demonstrate that Video-R1 achieves significant improvements on video reasoning benchmarks such as VideoMMMU and VSI-Bench, as well as on general video benchmarks including MVBench and TempCompass, etc. Notably, Video-R1-7B attains a 37.1% accuracy on video spatial reasoning benchmark VSI-bench, surpassing the commercial proprietary model GPT-4o. All code, models, and data are released in: https://github.com/tulerfeng/Video-R1.
title Video-R1: Reinforcing Video Reasoning in MLLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.21776