Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Wenhui, Song, Ruihua, Li, Jiaze, Ju, Jianzhong, Luo, Zhenbo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911380264714240
author Tan, Wenhui
Song, Ruihua
Li, Jiaze
Ju, Jianzhong
Luo, Zhenbo
author_facet Tan, Wenhui
Song, Ruihua
Li, Jiaze
Ju, Jianzhong
Luo, Zhenbo
contents Recent progress in multi-modal large language models (MLLMs) has significantly advanced video understanding. However, their performance on long-form videos remains limited by computational constraints and suboptimal frame selection. We present Think-Clip-Sample (TCS), a training-free framework that enhances long video understanding through two key components: (i) Multi-Query Reasoning, which generates multiple queries to capture complementary aspects of the question and video; and (ii) Clip-level Slow-Fast Sampling, which adaptively balances dense local details and sparse global context. Extensive experiments on MLVU, LongVideoBench, and VideoMME demonstrate that TCS consistently improves performance across different MLLMs, boosting up to 6.9% accuracy, and is capable of achieving comparable accuracy with 50% fewer inference time cost, highlighting both efficiency and efficacy of TCS on long video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11359
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
Tan, Wenhui
Song, Ruihua
Li, Jiaze
Ju, Jianzhong
Luo, Zhenbo
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent progress in multi-modal large language models (MLLMs) has significantly advanced video understanding. However, their performance on long-form videos remains limited by computational constraints and suboptimal frame selection. We present Think-Clip-Sample (TCS), a training-free framework that enhances long video understanding through two key components: (i) Multi-Query Reasoning, which generates multiple queries to capture complementary aspects of the question and video; and (ii) Clip-level Slow-Fast Sampling, which adaptively balances dense local details and sparse global context. Extensive experiments on MLVU, LongVideoBench, and VideoMME demonstrate that TCS consistently improves performance across different MLLMs, boosting up to 6.9% accuracy, and is capable of achieving comparable accuracy with 50% fewer inference time cost, highlighting both efficiency and efficacy of TCS on long video understanding.
title Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.11359