Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Hongyu, Han, Songhao, Liao, Yue, Luo, Junfeng, Gao, Jialin, Yan, Shuicheng, Liu, Si
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908389730156544
author Li, Hongyu
Han, Songhao
Liao, Yue
Luo, Junfeng
Gao, Jialin
Yan, Shuicheng
Liu, Si
author_facet Li, Hongyu
Han, Songhao
Liao, Yue
Luo, Junfeng
Gao, Jialin
Yan, Shuicheng
Liu, Si
contents Understanding real-world videos with complex semantics and long temporal dependencies remains a fundamental challenge in computer vision. Recent progress in multimodal large language models (MLLMs) has demonstrated strong capabilities in vision-language tasks, while reinforcement learning tuning (RLT) has further improved their reasoning abilities. In this work, we explore RLT as a post-training strategy to enhance the video-specific reasoning capabilities of MLLMs. Built upon the Group Relative Policy Optimization (GRPO) framework, we propose a dual-reward formulation that supervises both semantic and temporal reasoning through discrete and continuous reward signals. To facilitate effective preference-based optimization, we introduce a variance-aware data selection strategy based on repeated inference to identify samples that provide informative learning signals. We evaluate our approach across eight representative video understanding tasks, including VideoQA, Temporal Video Grounding, and Grounded VideoQA. Our method consistently outperforms supervised fine-tuning and existing RLT baselines, achieving superior performance with significantly less training data. These results underscore the importance of reward design and data selection in advancing reasoning-centric video understanding with MLLMs. Notably, The initial code release (two months ago) has now been expanded with updates, including optimized reward mechanisms and additional datasets. The latest version is available at https://github.com/appletea233/Temporal-R1 .
format Preprint
id arxiv_https___arxiv_org_abs_2506_01908
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
Li, Hongyu
Han, Songhao
Liao, Yue
Luo, Junfeng
Gao, Jialin
Yan, Shuicheng
Liu, Si
Computer Vision and Pattern Recognition
Understanding real-world videos with complex semantics and long temporal dependencies remains a fundamental challenge in computer vision. Recent progress in multimodal large language models (MLLMs) has demonstrated strong capabilities in vision-language tasks, while reinforcement learning tuning (RLT) has further improved their reasoning abilities. In this work, we explore RLT as a post-training strategy to enhance the video-specific reasoning capabilities of MLLMs. Built upon the Group Relative Policy Optimization (GRPO) framework, we propose a dual-reward formulation that supervises both semantic and temporal reasoning through discrete and continuous reward signals. To facilitate effective preference-based optimization, we introduce a variance-aware data selection strategy based on repeated inference to identify samples that provide informative learning signals. We evaluate our approach across eight representative video understanding tasks, including VideoQA, Temporal Video Grounding, and Grounded VideoQA. Our method consistently outperforms supervised fine-tuning and existing RLT baselines, achieving superior performance with significantly less training data. These results underscore the importance of reward design and data selection in advancing reasoning-centric video understanding with MLLMs. Notably, The initial code release (two months ago) has now been expanded with updates, including optimized reward mechanisms and additional datasets. The latest version is available at https://github.com/appletea233/Temporal-R1 .
title Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.01908