VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Jianxiang, Hong, Meisheng, Li, Jungang, Guo, Weiyu, Hu, Xuming, Xiong, Hui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911582014930944
author He, Jianxiang
Hong, Meisheng
Li, Jungang
Guo, Weiyu
Hu, Xuming
Xiong, Hui
author_facet He, Jianxiang
Hong, Meisheng
Li, Jungang
Guo, Weiyu
Hu, Xuming
Xiong, Hui
contents Multimodal large language models (MLLMs) demonstrate exceptional performance in vision-language tasks, yet their processing of long videos is constrained by input context length and high computational costs. Sparse frame sampling thus becomes a necessary preprocessing step, with sampled frame quality directly impacting downstream performance. Existing keyframe search algorithms achieve a balance between efficiency and sampled frame quality but heavily rely on the visual modality alone. This makes them difficult to adapt to text-related tasks and often leads to retrieval results deviating from core semantic content. To address this, we propose the VISUAL-SUBTITLE INTEGRATION (VSI), a multimodal keyframe retrieval framework. It employs a dual-branch collaborative retrieval approach combining Video Search and Subtitle Match to fuse complementary visual and textual information for precise localization. Experiments on LongVideoBench and VideoMME demonstrate that VSI achieves state-of-the-art accuracy in keyframe retrieval while delivering breakthrough performance in text-related tasks and exhibiting strong generalization across other tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06869
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
He, Jianxiang
Hong, Meisheng
Li, Jungang
Guo, Weiyu
Hu, Xuming
Xiong, Hui
Computer Vision and Pattern Recognition
Artificial Intelligence
I.2.10
Multimodal large language models (MLLMs) demonstrate exceptional performance in vision-language tasks, yet their processing of long videos is constrained by input context length and high computational costs. Sparse frame sampling thus becomes a necessary preprocessing step, with sampled frame quality directly impacting downstream performance. Existing keyframe search algorithms achieve a balance between efficiency and sampled frame quality but heavily rely on the visual modality alone. This makes them difficult to adapt to text-related tasks and often leads to retrieval results deviating from core semantic content. To address this, we propose the VISUAL-SUBTITLE INTEGRATION (VSI), a multimodal keyframe retrieval framework. It employs a dual-branch collaborative retrieval approach combining Video Search and Subtitle Match to fuse complementary visual and textual information for precise localization. Experiments on LongVideoBench and VideoMME demonstrate that VSI achieves state-of-the-art accuracy in keyframe retrieval while delivering breakthrough performance in text-related tasks and exhibiting strong generalization across other tasks.
title VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
I.2.10
url https://arxiv.org/abs/2508.06869