RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Malik, Sameer, Yamada, Moyuru, Singh, Ayush, Aggarwal, Dishank
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913822770462720
author Malik, Sameer
Yamada, Moyuru
Singh, Ayush
Aggarwal, Dishank
author_facet Malik, Sameer
Yamada, Moyuru
Singh, Ayush
Aggarwal, Dishank
contents Comprehending long videos remains a significant challenge for Large Multi-modal Models (LMMs). Current LMMs struggle to process even minutes to hours videos due to their lack of explicit memory and retrieval mechanisms. To address this limitation, we propose RAVU (Retrieval Augmented Video Understanding), a novel framework for video understanding enhanced by retrieval with compositional reasoning over a spatio-temporal graph. We construct a graph representation of the video, capturing both spatial and temporal relationships between entities. This graph serves as a long-term memory, allowing us to track objects and their actions across time. To answer complex queries, we decompose the queries into a sequence of reasoning steps and execute these steps on the graph, retrieving relevant key information. Our approach enables more accurate understanding of long videos, particularly for queries that require multi-hop reasoning and tracking objects across frames. Our approach demonstrate superior performances with limited retrieved frames (5-10) compared with other SOTA methods and baselines on two major video QA datasets, NExT-QA and EgoSchema.
format Preprint
id arxiv_https___arxiv_org_abs_2505_03173
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
Malik, Sameer
Yamada, Moyuru
Singh, Ayush
Aggarwal, Dishank
Computer Vision and Pattern Recognition
Artificial Intelligence
Comprehending long videos remains a significant challenge for Large Multi-modal Models (LMMs). Current LMMs struggle to process even minutes to hours videos due to their lack of explicit memory and retrieval mechanisms. To address this limitation, we propose RAVU (Retrieval Augmented Video Understanding), a novel framework for video understanding enhanced by retrieval with compositional reasoning over a spatio-temporal graph. We construct a graph representation of the video, capturing both spatial and temporal relationships between entities. This graph serves as a long-term memory, allowing us to track objects and their actions across time. To answer complex queries, we decompose the queries into a sequence of reasoning steps and execute these steps on the graph, retrieving relevant key information. Our approach enables more accurate understanding of long videos, particularly for queries that require multi-hop reasoning and tracking objects across frames. Our approach demonstrate superior performances with limited retrieved frames (5-10) compared with other SOTA methods and baselines on two major video QA datasets, NExT-QA and EgoSchema.
title RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.03173