Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Tao, Ju, Shaobo, Wu, Qiong, Fang, Chenxin, Zhang, Kun, Peng, Jun, Li, Hui, Zhou, Yiyi, Ji, Rongrong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914460372828160
author Chen, Tao
Ju, Shaobo
Wu, Qiong
Fang, Chenxin
Zhang, Kun
Peng, Jun
Li, Hui
Zhou, Yiyi
Ji, Rongrong
author_facet Chen, Tao
Ju, Shaobo
Wu, Qiong
Fang, Chenxin
Zhang, Kun
Peng, Jun
Li, Hui
Zhou, Yiyi
Ji, Rongrong
contents Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose an effective and efficient paradigm to remedy this shortcoming, termed One-shot video-Clip based Retrieval-Augmented Generation (OneClip-RAG). Compared with existing video RAG methods, OneClip-RAG makes full use of the merits of video clips for augmented video understanding in terms of both knowledge integrity and semantic coherence. Besides, it is also equipped with a novel query-guided video chunking algorithm that can unify clip chunking and cross-modal retrieval in one processing step, avoiding redundant computations. To improve instruction following, we further propose a new dataset called SynLongVideo and design a progressive training regime for OneClip-RAG. OneClip-RAG is plugged into three recent MLLMs and validated on a set of long-video benchmarks. Experimental results not only show the obvious performance gains by OneClip-RAG over MLLMs, e.g., boosting Qwen3-VL 8B to the level of GPT-5 on MLVU, but also show its superior efficiency in handling long videos. e.g., enabling LLaVA-Video understand up to an hour of videos in less than 1.2 minutes on a single 4090 GPU.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08410
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
Chen, Tao
Ju, Shaobo
Wu, Qiong
Fang, Chenxin
Zhang, Kun
Peng, Jun
Li, Hui
Zhou, Yiyi
Ji, Rongrong
Computer Vision and Pattern Recognition
Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose an effective and efficient paradigm to remedy this shortcoming, termed One-shot video-Clip based Retrieval-Augmented Generation (OneClip-RAG). Compared with existing video RAG methods, OneClip-RAG makes full use of the merits of video clips for augmented video understanding in terms of both knowledge integrity and semantic coherence. Besides, it is also equipped with a novel query-guided video chunking algorithm that can unify clip chunking and cross-modal retrieval in one processing step, avoiding redundant computations. To improve instruction following, we further propose a new dataset called SynLongVideo and design a progressive training regime for OneClip-RAG. OneClip-RAG is plugged into three recent MLLMs and validated on a set of long-video benchmarks. Experimental results not only show the obvious performance gains by OneClip-RAG over MLLMs, e.g., boosting Qwen3-VL 8B to the level of GPT-5 on MLVU, but also show its superior efficiency in handling long videos. e.g., enabling LLaVA-Video understand up to an hour of videos in less than 1.2 minutes on a single 4090 GPU.
title Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.08410