IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Yujia, Jiao, Jile, Feng, Xuetao, Ye, Zixuan, Wang, Yuan, Wang, Zhicheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911044188766208
author Liang, Yujia
Jiao, Jile
Feng, Xuetao
Ye, Zixuan
Wang, Yuan
Wang, Zhicheng
author_facet Liang, Yujia
Jiao, Jile
Feng, Xuetao
Ye, Zixuan
Wang, Yuan
Wang, Zhicheng
contents Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with varying camera angles or scene changes. This challenge can render failures such as instance identity forgetting and key frame negligence. In this work, we first attribute the challenge to the lack of multi-shot annotations among existing datasets and therefore we introduce a new dataset termed MultiClip-Bench, featuring dense descriptions and instruction-based question-answering pairs tailored for multi-shot scenarios. We empirically find that the training set significantly boosts the multi-shot performance, while the testing benchmark provides a reliable measure of the model capability in multi-shot scenarios. By further analyzing and discovering that current models only encode instance features in a discrete or lossy manner, at the risk of missing identity information, we then contribute a new model IPFormer-VideoLLM. Its key idea is the injection of instance-level features as instance prompts through an efficient attention-based connector. This allows for the aggregation of instance-specific information across scenes. Experiments demonstrate that our proposed dataset and model not only enhance the multi-scene video understanding significantly, but also offer distinct advantages across various video benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21116
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
Liang, Yujia
Jiao, Jile
Feng, Xuetao
Ye, Zixuan
Wang, Yuan
Wang, Zhicheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Video Large Language Models (VideoLLMs) have demonstrated remarkable understanding capabilities, but are found struggling to tackle multi-shot scenarios,e.g., video clips with varying camera angles or scene changes. This challenge can render failures such as instance identity forgetting and key frame negligence. In this work, we first attribute the challenge to the lack of multi-shot annotations among existing datasets and therefore we introduce a new dataset termed MultiClip-Bench, featuring dense descriptions and instruction-based question-answering pairs tailored for multi-shot scenarios. We empirically find that the training set significantly boosts the multi-shot performance, while the testing benchmark provides a reliable measure of the model capability in multi-shot scenarios. By further analyzing and discovering that current models only encode instance features in a discrete or lossy manner, at the risk of missing identity information, we then contribute a new model IPFormer-VideoLLM. Its key idea is the injection of instance-level features as instance prompts through an efficient attention-based connector. This allows for the aggregation of instance-specific information across scenes. Experiments demonstrate that our proposed dataset and model not only enhance the multi-scene video understanding significantly, but also offer distinct advantages across various video benchmarks.
title IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.21116