KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Baiyang, Peng, Jun, Zhang, Yuxin, Chen, Guangyao, Yang, Feidiao, Guo, Jianyuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912873343614976
author Song, Baiyang
Peng, Jun
Zhang, Yuxin
Chen, Guangyao
Yang, Feidiao
Guo, Jianyuan
author_facet Song, Baiyang
Peng, Jun
Zhang, Yuxin
Chen, Guangyao
Yang, Feidiao
Guo, Jianyuan
contents Training-free video understanding leverages the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating a video as a sequence of static frames, thus obviating the need for costly video-specific training. However, this paradigm often suffers from severe visual redundancy and high computational overhead, especially when processing long videos. Crucially, existing keyframe selection strategies, especially those based on CLIP similarity, are prone to biases and may inadvertently overlook critical frames, resulting in suboptimal video comprehension. To address these significant challenges, we propose \textbf{KTV}, a novel two-stage framework for efficient and effective training-free video understanding. In the first stage, KTV performs question-agnostic keyframe selection by clustering frame-level visual features, yielding a compact, diverse, and representative subset of frames that mitigates temporal redundancy. In the second stage, KTV applies key visual token selection, pruning redundant or less informative tokens from each selected keyframe based on token importance and redundancy, which significantly reduces the number of tokens fed into the LLM. Extensive experiments on the Multiple-Choice VideoQA task demonstrate that KTV outperforms state-of-the-art training-free baselines while using significantly fewer visual tokens, \emph{e.g.}, only 504 visual tokens for a 60-min video with 10800 frames, achieving $44.8\%$ accuracy on the MLVU-Test benchmark. In particular, KTV also exceeds several training-based approaches on certain benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03615
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
Song, Baiyang
Peng, Jun
Zhang, Yuxin
Chen, Guangyao
Yang, Feidiao
Guo, Jianyuan
Computer Vision and Pattern Recognition
Training-free video understanding leverages the strong image comprehension capabilities of pre-trained vision language models (VLMs) by treating a video as a sequence of static frames, thus obviating the need for costly video-specific training. However, this paradigm often suffers from severe visual redundancy and high computational overhead, especially when processing long videos. Crucially, existing keyframe selection strategies, especially those based on CLIP similarity, are prone to biases and may inadvertently overlook critical frames, resulting in suboptimal video comprehension. To address these significant challenges, we propose \textbf{KTV}, a novel two-stage framework for efficient and effective training-free video understanding. In the first stage, KTV performs question-agnostic keyframe selection by clustering frame-level visual features, yielding a compact, diverse, and representative subset of frames that mitigates temporal redundancy. In the second stage, KTV applies key visual token selection, pruning redundant or less informative tokens from each selected keyframe based on token importance and redundancy, which significantly reduces the number of tokens fed into the LLM. Extensive experiments on the Multiple-Choice VideoQA task demonstrate that KTV outperforms state-of-the-art training-free baselines while using significantly fewer visual tokens, \emph{e.g.}, only 504 visual tokens for a 60-min video with 10800 frames, achieving $44.8\%$ accuracy on the MLVU-Test benchmark. In particular, KTV also exceeds several training-based approaches on certain benchmarks.
title KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.03615