Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kwon, Minchan, Shon, Hyounguk, Kim, Junmo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911518835081216
author Kwon, Minchan
Shon, Hyounguk
Kim, Junmo
author_facet Kwon, Minchan
Shon, Hyounguk
Kim, Junmo
contents Large multimodal models (LMMs) have recently demonstrated remarkable performance in video question answering (VideoQA), yet reasoning over video remains challenging due to high inference cost and diluted information. Keyframe selection offers efficiency and sharper reasoning but suffers from sparse supervision and redundant frame choices when relying only on image-text similarity. We present a question-aware keyframe selection framework with two components: pseudo keyframe labels derived from LMMs that provide informative supervision and a coverage regularization that promotes diverse, complementary evidence across time. Experiments on NExT-QA show that our method significantly improves accuracy, especially for temporal and causal question types, establishing keyframe selection as an effective and learnable module for VideoQA.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14953
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
Kwon, Minchan
Shon, Hyounguk
Kim, Junmo
Computer Vision and Pattern Recognition
Artificial Intelligence
Large multimodal models (LMMs) have recently demonstrated remarkable performance in video question answering (VideoQA), yet reasoning over video remains challenging due to high inference cost and diluted information. Keyframe selection offers efficiency and sharper reasoning but suffers from sparse supervision and redundant frame choices when relying only on image-text similarity. We present a question-aware keyframe selection framework with two components: pseudo keyframe labels derived from LMMs that provide informative supervision and a coverage regularization that promotes diverse, complementary evidence across time. Experiments on NExT-QA show that our method significantly improves accuracy, especially for temporal and causal question types, establishing keyframe selection as an effective and learnable module for VideoQA.
title Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.14953