Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yiheng, Zhu, Lichen, Lin, Yueqian, Liu, Yudong, Zhang, Jingyang, Li, Hai "Helen", Chen, Yiran
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911560123809792
author Wang, Yiheng
Zhu, Lichen
Lin, Yueqian
Liu, Yudong
Zhang, Jingyang
Li, Hai "Helen"
Chen, Yiran
author_facet Wang, Yiheng
Zhu, Lichen
Lin, Yueqian
Liu, Yudong
Zhang, Jingyang
Li, Hai "Helen"
Chen, Yiran
contents Multimodal Large Language Models (MLLMs) have shown strong performance on video question answering, but their application to long-form videos is constrained by limited context length and computational cost, making keyframe sampling essential. Existing approaches typically rely on semantic relevance or reinforcement learning, which either fail to capture evidential clues or suffer from inefficient combinatorial optimization. In this work, we propose an evidence-driven keyframe sampling framework grounded in information bottleneck theory. We formulate keyframe selection as maximizing the conditional mutual information between selected frames and the query, providing a principled objective that reflects each frame's contribution to answering the question. To make this objective tractable, we exploit its structure to derive a decomposed optimization that reduces subset selection to independent frame-level scoring. We further introduce a query-conditioned evidence scoring network trained with a contrastive objective to estimate evidential importance efficiently. Experiments on long-form video understanding benchmarks show that our method consistently outperforms prior sampling strategies under strict token budgets, while significantly improving training efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01002
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
Wang, Yiheng
Zhu, Lichen
Lin, Yueqian
Liu, Yudong
Zhang, Jingyang
Li, Hai "Helen"
Chen, Yiran
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimodal Large Language Models (MLLMs) have shown strong performance on video question answering, but their application to long-form videos is constrained by limited context length and computational cost, making keyframe sampling essential. Existing approaches typically rely on semantic relevance or reinforcement learning, which either fail to capture evidential clues or suffer from inefficient combinatorial optimization. In this work, we propose an evidence-driven keyframe sampling framework grounded in information bottleneck theory. We formulate keyframe selection as maximizing the conditional mutual information between selected frames and the query, providing a principled objective that reflects each frame's contribution to answering the question. To make this objective tractable, we exploit its structure to derive a decomposed optimization that reduces subset selection to independent frame-level scoring. We further introduce a query-conditioned evidence scoring network trained with a contrastive objective to estimate evidential importance efficiently. Experiments on long-form video understanding benchmarks show that our method consistently outperforms prior sampling strategies under strict token budgets, while significantly improving training efficiency.
title Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2604.01002