FrameOracle: Learning What to See and How Much to See in Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914292688748544 |
|---|---|
| author | Li, Chaoyu Li, Tianzhi Tao, Fei Zhao, Zhenyu Wu, Ziqian Zhao, Maozheng Song, Juntong Niu, Cheng Fazli, Pooyan |
| author_facet | Li, Chaoyu Li, Tianzhi Tao, Fei Zhao, Zhenyu Wu, Ziqian Zhao, Maozheng Song, Juntong Niu, Cheng Fazli, Pooyan |
| contents | Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of frames. Existing frame sampling strategies, such as uniform or fixed-budget selection, fail to adapt to variations in content density or task complexity. To address this, we present FrameOracle, a lightweight, plug-and-play module that predicts both (1) which frames are most relevant to a given query and (2) how many frames are needed. FrameOracle is trained via a curriculum that progresses from weak proxy signals, such as cross-modal similarity, to stronger supervision with FrameOracle-41K, the first large-scale VideoQA dataset with validated keyframe annotations specifying minimal sufficient frames per question. Extensive experiments across five VLMs and six benchmarks show that FrameOracle reduces 16-frame inputs to an average of 10.4 frames without accuracy loss. When starting from 64-frame candidates, it reduces inputs to 13.9 frames on average while improving accuracy by 1.5%, achieving state-of-the-art efficiency-accuracy trade-offs for scalable video understanding. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_03584 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FrameOracle: Learning What to See and How Much to See in Videos Li, Chaoyu Li, Tianzhi Tao, Fei Zhao, Zhenyu Wu, Ziqian Zhao, Maozheng Song, Juntong Niu, Cheng Fazli, Pooyan Computer Vision and Pattern Recognition Vision-language models (VLMs) advance video understanding but operate under tight computational budgets, making performance dependent on selecting a small, high-quality subset of frames. Existing frame sampling strategies, such as uniform or fixed-budget selection, fail to adapt to variations in content density or task complexity. To address this, we present FrameOracle, a lightweight, plug-and-play module that predicts both (1) which frames are most relevant to a given query and (2) how many frames are needed. FrameOracle is trained via a curriculum that progresses from weak proxy signals, such as cross-modal similarity, to stronger supervision with FrameOracle-41K, the first large-scale VideoQA dataset with validated keyframe annotations specifying minimal sufficient frames per question. Extensive experiments across five VLMs and six benchmarks show that FrameOracle reduces 16-frame inputs to an average of 10.4 frames without accuracy loss. When starting from 64-frame candidates, it reduces inputs to 13.9 frames on average while improving accuracy by 1.5%, achieving state-of-the-art efficiency-accuracy trade-offs for scalable video understanding. |
| title | FrameOracle: Learning What to See and How Much to See in Videos |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2510.03584 |