VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911582014930944 |
|---|---|
| author | He, Jianxiang Hong, Meisheng Li, Jungang Guo, Weiyu Hu, Xuming Xiong, Hui |
| author_facet | He, Jianxiang Hong, Meisheng Li, Jungang Guo, Weiyu Hu, Xuming Xiong, Hui |
| contents | Multimodal large language models (MLLMs) demonstrate exceptional performance in vision-language tasks, yet their processing of long videos is constrained by input context length and high computational costs. Sparse frame sampling thus becomes a necessary preprocessing step, with sampled frame quality directly impacting downstream performance. Existing keyframe search algorithms achieve a balance between efficiency and sampled frame quality but heavily rely on the visual modality alone. This makes them difficult to adapt to text-related tasks and often leads to retrieval results deviating from core semantic content. To address this, we propose the VISUAL-SUBTITLE INTEGRATION (VSI), a multimodal keyframe retrieval framework. It employs a dual-branch collaborative retrieval approach combining Video Search and Subtitle Match to fuse complementary visual and textual information for precise localization. Experiments on LongVideoBench and VideoMME demonstrate that VSI achieves state-of-the-art accuracy in keyframe retrieval while delivering breakthrough performance in text-related tasks and exhibiting strong generalization across other tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_06869 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding He, Jianxiang Hong, Meisheng Li, Jungang Guo, Weiyu Hu, Xuming Xiong, Hui Computer Vision and Pattern Recognition Artificial Intelligence I.2.10 Multimodal large language models (MLLMs) demonstrate exceptional performance in vision-language tasks, yet their processing of long videos is constrained by input context length and high computational costs. Sparse frame sampling thus becomes a necessary preprocessing step, with sampled frame quality directly impacting downstream performance. Existing keyframe search algorithms achieve a balance between efficiency and sampled frame quality but heavily rely on the visual modality alone. This makes them difficult to adapt to text-related tasks and often leads to retrieval results deviating from core semantic content. To address this, we propose the VISUAL-SUBTITLE INTEGRATION (VSI), a multimodal keyframe retrieval framework. It employs a dual-branch collaborative retrieval approach combining Video Search and Subtitle Match to fuse complementary visual and textual information for precise localization. Experiments on LongVideoBench and VideoMME demonstrate that VSI achieves state-of-the-art accuracy in keyframe retrieval while delivering breakthrough performance in text-related tasks and exhibiting strong generalization across other tasks. |
| title | VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence I.2.10 |
| url | https://arxiv.org/abs/2508.06869 |