Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qu, Tianyuan, Tang, Longxiang, Peng, Bohao, Yang, Senqiao, Yu, Bei, Jia, Jiaya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916664422957056
author Qu, Tianyuan
Tang, Longxiang
Peng, Bohao
Yang, Senqiao
Yu, Bei
Jia, Jiaya
author_facet Qu, Tianyuan
Tang, Longxiang
Peng, Bohao
Yang, Senqiao
Yu, Bei
Jia, Jiaya
contents The rise of Large Vision-Language Models (LVLMs) has significantly advanced video understanding. However, efficiently processing long videos remains a challenge due to the ``Sampling Dilemma'': low-density sampling risks missing critical information, while high-density sampling introduces redundancy. To address this issue, we introduce LSDBench, the first benchmark designed to evaluate LVLMs on long-video tasks by constructing high Necessary Sampling Density (NSD) questions, where NSD represents the minimum sampling density required to accurately answer a given question. LSDBench focuses on dense, short-duration actions to rigorously assess the sampling strategies employed by LVLMs. To tackle the challenges posed by high-NSD questions, we propose a novel Reasoning-Driven Hierarchical Sampling (RHS) framework, which combines global localization of question-relevant cues with local dense sampling for precise inference. Additionally, we develop a lightweight Semantic-Guided Frame Selector to prioritize informative frames, enabling RHS to achieve comparable or superior performance with significantly fewer sampled frames. Together, our LSDBench and RHS framework address the unique challenges of high-NSD long-video tasks, setting a new standard for evaluating and improving LVLMs in this domain. Our benchmark and evaluation codes has been released at: https://github.com/dvlab-research/LSDBench
format Preprint
id arxiv_https___arxiv_org_abs_2503_12496
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
Qu, Tianyuan
Tang, Longxiang
Peng, Bohao
Yang, Senqiao
Yu, Bei
Jia, Jiaya
Computer Vision and Pattern Recognition
The rise of Large Vision-Language Models (LVLMs) has significantly advanced video understanding. However, efficiently processing long videos remains a challenge due to the ``Sampling Dilemma'': low-density sampling risks missing critical information, while high-density sampling introduces redundancy. To address this issue, we introduce LSDBench, the first benchmark designed to evaluate LVLMs on long-video tasks by constructing high Necessary Sampling Density (NSD) questions, where NSD represents the minimum sampling density required to accurately answer a given question. LSDBench focuses on dense, short-duration actions to rigorously assess the sampling strategies employed by LVLMs. To tackle the challenges posed by high-NSD questions, we propose a novel Reasoning-Driven Hierarchical Sampling (RHS) framework, which combines global localization of question-relevant cues with local dense sampling for precise inference. Additionally, we develop a lightweight Semantic-Guided Frame Selector to prioritize informative frames, enabling RHS to achieve comparable or superior performance with significantly fewer sampled frames. Together, our LSDBench and RHS framework address the unique challenges of high-NSD long-video tasks, setting a new standard for evaluating and improving LVLMs in this domain. Our benchmark and evaluation codes has been released at: https://github.com/dvlab-research/LSDBench
title Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.12496