Moment Sampling in Video LLMs for Long-Form Video QA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chasmai, Mustafa, Jagatap, Gauri, KV, Gouthaman, Van Horn, Grant, Maji, Subhransu, Fanelli, Andrea
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913919835045888
author Chasmai, Mustafa
Jagatap, Gauri
KV, Gouthaman
Van Horn, Grant
Maji, Subhransu
Fanelli, Andrea
author_facet Chasmai, Mustafa
Jagatap, Gauri
KV, Gouthaman
Van Horn, Grant
Maji, Subhransu
Fanelli, Andrea
contents Recent advancements in video large language models (Video LLMs) have significantly advanced the field of video question answering (VideoQA). While existing methods perform well on short videos, they often struggle with long-range reasoning in longer videos. To scale Video LLMs for longer video content, frame sub-sampling (selecting frames at regular intervals) is commonly used. However, this approach is suboptimal, often leading to the loss of crucial frames or the inclusion of redundant information from multiple similar frames. Missing key frames impairs the model's ability to answer questions accurately, while redundant frames lead the model to focus on irrelevant video segments and increase computational resource consumption. In this paper, we investigate the use of a general-purpose text-to-video moment retrieval model to guide the frame sampling process. We propose "moment sampling", a novel, model-agnostic approach that enables the model to select the most relevant frames according to the context of the question. Specifically, we employ a lightweight moment retrieval model to prioritize frame selection. By focusing on the frames most pertinent to the given question, our method enhances long-form VideoQA performance in Video LLMs. Through extensive experiments on four long-form VideoQA datasets, using four state-of-the-art Video LLMs, we demonstrate the effectiveness of the proposed approach.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00033
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Moment Sampling in Video LLMs for Long-Form Video QA
Chasmai, Mustafa
Jagatap, Gauri
KV, Gouthaman
Van Horn, Grant
Maji, Subhransu
Fanelli, Andrea
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Recent advancements in video large language models (Video LLMs) have significantly advanced the field of video question answering (VideoQA). While existing methods perform well on short videos, they often struggle with long-range reasoning in longer videos. To scale Video LLMs for longer video content, frame sub-sampling (selecting frames at regular intervals) is commonly used. However, this approach is suboptimal, often leading to the loss of crucial frames or the inclusion of redundant information from multiple similar frames. Missing key frames impairs the model's ability to answer questions accurately, while redundant frames lead the model to focus on irrelevant video segments and increase computational resource consumption. In this paper, we investigate the use of a general-purpose text-to-video moment retrieval model to guide the frame sampling process. We propose "moment sampling", a novel, model-agnostic approach that enables the model to select the most relevant frames according to the context of the question. Specifically, we employ a lightweight moment retrieval model to prioritize frame selection. By focusing on the frames most pertinent to the given question, our method enhances long-form VideoQA performance in Video LLMs. Through extensive experiments on four long-form VideoQA datasets, using four state-of-the-art Video LLMs, we demonstrate the effectiveness of the proposed approach.
title Moment Sampling in Video LLMs for Long-Form Video QA
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.00033