MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Xiaokun, Wu, Zezhong, Ding, Zewen, Xu, Linli
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917187971710976
author Sun, Xiaokun
Wu, Zezhong
Ding, Zewen
Xu, Linli
author_facet Sun, Xiaokun
Wu, Zezhong
Ding, Zewen
Xu, Linli
contents Reinforcement learning based post-training paradigms for Video Large Language Models (VideoLLMs) have achieved significant success by optimizing for visual-semantic tasks such as captioning or VideoQA. However, while these approaches effectively enhance perception abilities, they primarily target holistic content understanding, often lacking explicit supervision for intrinsic temporal coherence and inter-frame correlations. This tendency limits the models' ability to capture intricate dynamics and fine-grained visual causality. To explicitly bridge this gap, we propose a novel post-training objective: Masked Video Prediction (MVP). By requiring the model to reconstruct a masked continuous segment from a set of challenging distractors, MVP forces the model to attend to the sequential logic and temporal context of events. To support scalable training, we introduce a scalable data synthesis pipeline capable of transforming arbitrary video corpora into MVP training samples, and further employ Group Relative Policy Optimization (GRPO) with a fine-grained reward function to enhance the model's understanding of video context and temporal properties. Comprehensive evaluations demonstrate that MVP enhances video reasoning capabilities by directly reinforcing temporal reasoning and causal understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03781
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
Sun, Xiaokun
Wu, Zezhong
Ding, Zewen
Xu, Linli
Computer Vision and Pattern Recognition
Reinforcement learning based post-training paradigms for Video Large Language Models (VideoLLMs) have achieved significant success by optimizing for visual-semantic tasks such as captioning or VideoQA. However, while these approaches effectively enhance perception abilities, they primarily target holistic content understanding, often lacking explicit supervision for intrinsic temporal coherence and inter-frame correlations. This tendency limits the models' ability to capture intricate dynamics and fine-grained visual causality. To explicitly bridge this gap, we propose a novel post-training objective: Masked Video Prediction (MVP). By requiring the model to reconstruct a masked continuous segment from a set of challenging distractors, MVP forces the model to attend to the sequential logic and temporal context of events. To support scalable training, we introduce a scalable data synthesis pipeline capable of transforming arbitrary video corpora into MVP training samples, and further employ Group Relative Policy Optimization (GRPO) with a fine-grained reward function to enhance the model's understanding of video context and temporal properties. Comprehensive evaluations demonstrate that MVP enhances video reasoning capabilities by directly reinforcing temporal reasoning and causal understanding.
title MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.03781