FrameOracle: Learning What to See and How Much to See in Videos

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Chaoyu, Li, Tianzhi, Tao, Fei, Zhao, Zhenyu, Wu, Ziqian, Zhao, Maozheng, Song, Juntong, Niu, Cheng, Fazli, Pooyan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914292688748544
author Li, Chaoyu
Li, Tianzhi
Tao, Fei
Zhao, Zhenyu
Wu, Ziqian
Zhao, Maozheng
Song, Juntong
Niu, Cheng
Fazli, Pooyan
author_facet Li, Chaoyu
Li, Tianzhi
Tao, Fei
Zhao, Zhenyu
Wu, Ziqian
Zhao, Maozheng
Song, Juntong
Niu, Cheng
Fazli, Pooyan
contents Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of frames. Existing frame sampling strategies, such as uniform or fixed-budget selection, fail to adapt to variations in content density or task complexity. To address this, we present FrameOracle, a lightweight, plug-and-play module that predicts both (1) which frames are most relevant to a given query and (2) how many frames are needed. FrameOracle is trained via a curriculum that progresses from weak proxy signals, such as cross-modal similarity, to stronger supervision with FrameOracle-41K, the first large-scale VideoQA dataset with validated keyframe annotations specifying minimal sufficient frames per question. Extensive experiments across five VLMs and six benchmarks show that FrameOracle reduces 16-frame inputs to an average of 10.4 frames without accuracy loss. When starting from 64-frame candidates, it reduces inputs to 13.9 frames on average while improving accuracy by 1.5%, achieving state-of-the-art efficiency-accuracy trade-offs for scalable video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03584
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FrameOracle: Learning What to See and How Much to See in Videos
Li, Chaoyu
Li, Tianzhi
Tao, Fei
Zhao, Zhenyu
Wu, Ziqian
Zhao, Maozheng
Song, Juntong
Niu, Cheng
Fazli, Pooyan
Computer Vision and Pattern Recognition
Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of frames. Existing frame sampling strategies, such as uniform or fixed-budget selection, fail to adapt to variations in content density or task complexity. To address this, we present FrameOracle, a lightweight, plug-and-play module that predicts both (1) which frames are most relevant to a given query and (2) how many frames are needed. FrameOracle is trained via a curriculum that progresses from weak proxy signals, such as cross-modal similarity, to stronger supervision with FrameOracle-41K, the first large-scale VideoQA dataset with validated keyframe annotations specifying minimal sufficient frames per question. Extensive experiments across five VLMs and six benchmarks show that FrameOracle reduces 16-frame inputs to an average of 10.4 frames without accuracy loss. When starting from 64-frame candidates, it reduces inputs to 13.9 frames on average while improving accuracy by 1.5%, achieving state-of-the-art efficiency-accuracy trade-offs for scalable video understanding.
title FrameOracle: Learning What to See and How Much to See in Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.03584