VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Honghao, Xu, Miao, Wang, Yiwei, Zhang, Dailing, Liu, Jun, Cai, Yujun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913037946978304
author Fu, Honghao
Xu, Miao
Wang, Yiwei
Zhang, Dailing
Liu, Jun
Cai, Yujun
author_facet Fu, Honghao
Xu, Miao
Wang, Yiwei
Zhang, Dailing
Liu, Jun
Cai, Yujun
contents Scaling multimodal large language models (MLLMs) to long videos is constrained by limited context windows. While retrieval-augmented generation (RAG) is a promising remedy by organizing query-relevant visual evidence into a compact context, most existing methods (i) flatten videos into independent segments, breaking their inherent spatio-temporal structure, and (ii) depend on explicit semantic matching, which can miss cues that are implicitly relevant to the query's intent. To overcome these limitations, we propose VideoStir, a structured and intent-aware long-video RAG framework. It firstly structures a video as a spatio-temporal graph at clip level, and then performs multi-hop retrieval to aggregate evidence across distant yet contextually related events. Furthermore, it introduces an MLLM-backed intent-relevance scorer that retrieves frames based on their alignment with the query's reasoning intent. To support this capability, we curate IR-600K, a large-scale dataset tailored for learning frame-query intent alignment. Experiments show that VideoStir is competitive with state-of-the-art baselines without relying on auxiliary information, highlighting the promise of shifting long-video RAG from flattened semantic matching to structured, intent-aware reasoning. Codes and checkpoints are available at https://github.com/RomGai/VideoStir.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05418
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
Fu, Honghao
Xu, Miao
Wang, Yiwei
Zhang, Dailing
Liu, Jun
Cai, Yujun
Computer Vision and Pattern Recognition
Artificial Intelligence
Scaling multimodal large language models (MLLMs) to long videos is constrained by limited context windows. While retrieval-augmented generation (RAG) is a promising remedy by organizing query-relevant visual evidence into a compact context, most existing methods (i) flatten videos into independent segments, breaking their inherent spatio-temporal structure, and (ii) depend on explicit semantic matching, which can miss cues that are implicitly relevant to the query's intent. To overcome these limitations, we propose VideoStir, a structured and intent-aware long-video RAG framework. It firstly structures a video as a spatio-temporal graph at clip level, and then performs multi-hop retrieval to aggregate evidence across distant yet contextually related events. Furthermore, it introduces an MLLM-backed intent-relevance scorer that retrieves frames based on their alignment with the query's reasoning intent. To support this capability, we curate IR-600K, a large-scale dataset tailored for learning frame-query intent alignment. Experiments show that VideoStir is competitive with state-of-the-art baselines without relying on auxiliary information, highlighting the promise of shifting long-video RAG from flattened semantic matching to structured, intent-aware reasoning. Codes and checkpoints are available at https://github.com/RomGai/VideoStir.
title VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2604.05418