M-LLM Based Video Frame Selection for Efficient Video Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hu, Kai, Gao, Feng, Nie, Xiaohan, Zhou, Peng, Tran, Son, Neiman, Tal, Wang, Lingyun, Shah, Mubarak, Hamid, Raffay, Yin, Bing, Chilimbi, Trishul
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912296152858624
author Hu, Kai
Gao, Feng
Nie, Xiaohan
Zhou, Peng
Tran, Son
Neiman, Tal
Wang, Lingyun
Shah, Mubarak
Hamid, Raffay
Yin, Bing
Chilimbi, Trishul
author_facet Hu, Kai
Gao, Feng
Nie, Xiaohan
Zhou, Peng
Tran, Son
Neiman, Tal
Wang, Lingyun
Shah, Mubarak
Hamid, Raffay
Yin, Bing
Chilimbi, Trishul
contents Recent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply naive uniform sampling to reduce the number of video frames that are fed into an M-LLM, particularly for long context videos. However, it could lose crucial context in certain periods of a video, so that the downstream M-LLM may not have sufficient visual information to answer a question. To attack this pain point, we propose a light-weight M-LLM -based frame selection method that adaptively select frames that are more relevant to users' queries. In order to train the proposed frame selector, we introduce two supervision signals (i) Spatial signal, where single frame importance score by prompting a M-LLM; (ii) Temporal signal, in which multiple frames selection by prompting Large Language Model (LLM) using the captions of all frame candidates. The selected frames are then digested by a frozen downstream video M-LLM for visual reasoning and question answering. Empirical results show that the proposed M-LLM video frame selector improves the performances various downstream video Large Language Model (video-LLM) across medium (ActivityNet, NExT-QA) and long (EgoSchema, LongVideoBench) context video question answering benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19680
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle M-LLM Based Video Frame Selection for Efficient Video Understanding
Hu, Kai
Gao, Feng
Nie, Xiaohan
Zhou, Peng
Tran, Son
Neiman, Tal
Wang, Lingyun
Shah, Mubarak
Hamid, Raffay
Yin, Bing
Chilimbi, Trishul
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply naive uniform sampling to reduce the number of video frames that are fed into an M-LLM, particularly for long context videos. However, it could lose crucial context in certain periods of a video, so that the downstream M-LLM may not have sufficient visual information to answer a question. To attack this pain point, we propose a light-weight M-LLM -based frame selection method that adaptively select frames that are more relevant to users' queries. In order to train the proposed frame selector, we introduce two supervision signals (i) Spatial signal, where single frame importance score by prompting a M-LLM; (ii) Temporal signal, in which multiple frames selection by prompting Large Language Model (LLM) using the captions of all frame candidates. The selected frames are then digested by a frozen downstream video M-LLM for visual reasoning and question answering. Empirical results show that the proposed M-LLM video frame selector improves the performances various downstream video Large Language Model (video-LLM) across medium (ActivityNet, NExT-QA) and long (EgoSchema, LongVideoBench) context video question answering benchmarks.
title M-LLM Based Video Frame Selection for Efficient Video Understanding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2502.19680