VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Fufangchen, Zhang, Liao, Shi, Daiqi, Gao, Yuanjun, Ye, Chen, Cai, Yang, Gao, Jian, Yan, Danfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909920436158464
author Zhao, Fufangchen
Zhang, Liao
Shi, Daiqi
Gao, Yuanjun
Ye, Chen
Cai, Yang
Gao, Jian
Yan, Danfeng
author_facet Zhao, Fufangchen
Zhang, Liao
Shi, Daiqi
Gao, Yuanjun
Ye, Chen
Cai, Yang
Gao, Jian
Yan, Danfeng
contents We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to reason about brief actions in short clips or rare transient events in long videos. VideoPerceiver adopts a two-stage training framework. During supervised fine-tuning (SFT), we construct "key-information-missing" videos by extracting event-action keywords from captions, identifying corresponding key frames, and replacing them with adjacent frames. We jointly encode original and modified video tokens with text tokens, aligning intermediate visual representations with keywords via an auxiliary contrastive loss to enhance sensitivity to fine-grained motion cues. In reinforcement learning (RL), both video variants are fed into the model to generate descriptions, and a novel relative reward ensures responses from complete videos outperform those from degraded inputs, explicitly training the model to recover temporally precise action details. We also curate a dataset of 80,000 videos with fine-grained actions and transient events. Experiments show VideoPerceiver substantially outperforms state-of-the-art VMLLMs on fine-grained action understanding and rare event captioning benchmarks, while maintaining strong performance on standard tasks. By prioritizing task-relevant visual features, our work redefines video-language model training for fine-grained perception.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18823
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
Zhao, Fufangchen
Zhang, Liao
Shi, Daiqi
Gao, Yuanjun
Ye, Chen
Cai, Yang
Gao, Jian
Yan, Danfeng
Computer Vision and Pattern Recognition
We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to reason about brief actions in short clips or rare transient events in long videos. VideoPerceiver adopts a two-stage training framework. During supervised fine-tuning (SFT), we construct "key-information-missing" videos by extracting event-action keywords from captions, identifying corresponding key frames, and replacing them with adjacent frames. We jointly encode original and modified video tokens with text tokens, aligning intermediate visual representations with keywords via an auxiliary contrastive loss to enhance sensitivity to fine-grained motion cues. In reinforcement learning (RL), both video variants are fed into the model to generate descriptions, and a novel relative reward ensures responses from complete videos outperform those from degraded inputs, explicitly training the model to recover temporally precise action details. We also curate a dataset of 80,000 videos with fine-grained actions and transient events. Experiments show VideoPerceiver substantially outperforms state-of-the-art VMLLMs on fine-grained action understanding and rare event captioning benchmarks, while maintaining strong performance on standard tasks. By prioritizing task-relevant visual features, our work redefines video-language model training for fine-grained perception.
title VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18823