Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Bo, Wu, Wenhao, Wu, Qiangqiang, Song, Yuxin, Chan, Antoni B.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912403768213504
author Fang, Bo
Wu, Wenhao
Wu, Qiangqiang
Song, Yuxin
Chan, Antoni B.
author_facet Fang, Bo
Wu, Wenhao
Wu, Qiangqiang
Song, Yuxin
Chan, Antoni B.
contents Employing Multimodal Large Language Models (MLLMs) for long video understanding remains a challenging problem due to the dilemma between the substantial number of video frames (i.e., visual tokens) versus the limited context length of language models. Traditional uniform sampling often leads to selection of irrelevant content, while post-training MLLMs on thousands of frames imposes a substantial computational burden. In this paper, we propose threading keyframes with narratives (Nar-KFC), a plug-and-play module to facilitate effective and efficient long video perception. Nar-KFC generally involves two collaborative steps. First, we formulate the keyframe selection process as an integer quadratic programming problem, jointly optimizing query-relevance and frame-diversity. To avoid its computational complexity, a customized greedy search strategy is designed as an efficient alternative. Second, to mitigate the temporal discontinuity caused by sparse keyframe sampling, we further introduce interleaved textual narratives generated from non-keyframes using off-the-shelf captioners. These narratives are inserted between keyframes based on their true temporal order, forming a coherent and compact representation. Nar-KFC thus serves as a temporal- and content-aware compression strategy that complements visual and textual modalities. Experimental results on multiple long-video benchmarks demonstrate that Nar-KFC significantly improves the performance of popular MLLMs. Code will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24158
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
Fang, Bo
Wu, Wenhao
Wu, Qiangqiang
Song, Yuxin
Chan, Antoni B.
Computer Vision and Pattern Recognition
Employing Multimodal Large Language Models (MLLMs) for long video understanding remains a challenging problem due to the dilemma between the substantial number of video frames (i.e., visual tokens) versus the limited context length of language models. Traditional uniform sampling often leads to selection of irrelevant content, while post-training MLLMs on thousands of frames imposes a substantial computational burden. In this paper, we propose threading keyframes with narratives (Nar-KFC), a plug-and-play module to facilitate effective and efficient long video perception. Nar-KFC generally involves two collaborative steps. First, we formulate the keyframe selection process as an integer quadratic programming problem, jointly optimizing query-relevance and frame-diversity. To avoid its computational complexity, a customized greedy search strategy is designed as an efficient alternative. Second, to mitigate the temporal discontinuity caused by sparse keyframe sampling, we further introduce interleaved textual narratives generated from non-keyframes using off-the-shelf captioners. These narratives are inserted between keyframes based on their true temporal order, forming a coherent and compact representation. Nar-KFC thus serves as a temporal- and content-aware compression strategy that complements visual and textual modalities. Experimental results on multiple long-video benchmarks demonstrate that Nar-KFC significantly improves the performance of popular MLLMs. Code will be made publicly available.
title Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.24158