RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
Fuente:
arXiv
Saved in:
| Main Authors: | Malik, Sameer, Yamada, Moyuru, Singh, Ayush, Aggarwal, Dishank |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GLoD: Composing Global Contexts and Local Details in Image Generation
by: Yamada, Moyuru
Published: (2024)
by: Yamada, Moyuru
Published: (2024)
HICO-DET-SG and V-COCO-SG: New Data Splits for Evaluating the Systematic Generalization Performance of Human-Object Interaction Detection Models
by: Takemoto, Kentaro, et al.
Published: (2023)
by: Takemoto, Kentaro, et al.
Published: (2023)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
by: Zeng, Nianbo, et al.
Published: (2025)
by: Zeng, Nianbo, et al.
Published: (2025)
NutriScreener: Retrieval-Augmented Multi-Pose Graph Attention Network for Malnourishment Screening
by: Khan, Misaal, et al.
Published: (2025)
by: Khan, Misaal, et al.
Published: (2025)
D3: Data Diversity Design for Systematic Generalization in Visual Question Answering
by: Rahimi, Amir, et al.
Published: (2023)
by: Rahimi, Amir, et al.
Published: (2023)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
by: Madasu, Avinash, et al.
Published: (2023)
by: Madasu, Avinash, et al.
Published: (2023)
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios
by: Yan, Peizheng, et al.
Published: (2026)
by: Yan, Peizheng, et al.
Published: (2026)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
by: Luo, Yongdong, et al.
Published: (2024)
by: Luo, Yongdong, et al.
Published: (2024)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
by: Wang, Zun, et al.
Published: (2024)
by: Wang, Zun, et al.
Published: (2024)
VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding
by: Chen, Houlun, et al.
Published: (2024)
by: Chen, Houlun, et al.
Published: (2024)
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
by: Hu, Lanxiang, et al.
Published: (2025)
by: Hu, Lanxiang, et al.
Published: (2025)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
VideoRAG: Retrieval-Augmented Generation over Video Corpus
by: Jeong, Soyeong, et al.
Published: (2025)
by: Jeong, Soyeong, et al.
Published: (2025)
Predictive Reasoning with Augmented Anomaly Contrastive Learning for Compositional Visual Relations
by: Li, Chengtai, et al.
Published: (2026)
by: Li, Chengtai, et al.
Published: (2026)
Enhancing Supervised Composed Image Retrieval via Reasoning-Augmented Representation Engineering
by: Li, Jun, et al.
Published: (2025)
by: Li, Jun, et al.
Published: (2025)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
by: Kim, Junho, et al.
Published: (2024)
by: Kim, Junho, et al.
Published: (2024)
MADTempo: An Interactive System for Multi-Event Temporal Video Retrieval with Query Augmentation
by: Vu, Huu-An, et al.
Published: (2025)
by: Vu, Huu-An, et al.
Published: (2025)
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
by: Saravanan, Darshana, et al.
Published: (2024)
by: Saravanan, Darshana, et al.
Published: (2024)
Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
by: Masala, Mihai, et al.
Published: (2025)
by: Masala, Mihai, et al.
Published: (2025)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
by: Huang, De-An, et al.
Published: (2025)
by: Huang, De-An, et al.
Published: (2025)
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
by: Kong, Fanqi, et al.
Published: (2025)
by: Kong, Fanqi, et al.
Published: (2025)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
by: Deng, Andong, et al.
Published: (2024)
by: Deng, Andong, et al.
Published: (2024)
Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation
by: Vatsa, Mayank, et al.
Published: (2025)
by: Vatsa, Mayank, et al.
Published: (2025)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
by: Shen, Yucheng, et al.
Published: (2026)
by: Shen, Yucheng, et al.
Published: (2026)
FastV-RAG: Towards Fast and Fine-Grained Video QA with Retrieval-Augmented Generation
by: Li, Gen, et al.
Published: (2026)
by: Li, Gen, et al.
Published: (2026)
RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
by: Zare, Ali, et al.
Published: (2024)
by: Zare, Ali, et al.
Published: (2024)
Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection
by: Yakun, Cui, et al.
Published: (2025)
by: Yakun, Cui, et al.
Published: (2025)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
by: Ren, Xubin, et al.
Published: (2025)
by: Ren, Xubin, et al.
Published: (2025)
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
by: Rosa, Kevin Dela
Published: (2024)
by: Rosa, Kevin Dela
Published: (2024)
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
by: Durante, Zane, et al.
Published: (2026)
by: Durante, Zane, et al.
Published: (2026)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
by: Wang, Xingrui, et al.
Published: (2024)
by: Wang, Xingrui, et al.
Published: (2024)
Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
by: Ming, Yifei, et al.
Published: (2024)
by: Ming, Yifei, et al.
Published: (2024)
mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA
by: Yuan, Xu, et al.
Published: (2025)
by: Yuan, Xu, et al.
Published: (2025)
TurtleBench: A Visual Programming Benchmark in Turtle Geometry
by: Rismanchian, Sina, et al.
Published: (2024)
by: Rismanchian, Sina, et al.
Published: (2024)
Retrieval-Augmented Prompt for OOD Detection
by: Han, Ruisong, et al.
Published: (2025)
by: Han, Ruisong, et al.
Published: (2025)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
Similar Items
-
GLoD: Composing Global Contexts and Local Details in Image Generation
by: Yamada, Moyuru
Published: (2024) -
HICO-DET-SG and V-COCO-SG: New Data Splits for Evaluating the Systematic Generalization Performance of Human-Object Interaction Detection Models
by: Takemoto, Kentaro, et al.
Published: (2023) -
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
by: Zeng, Nianbo, et al.
Published: (2025) -
NutriScreener: Retrieval-Augmented Multi-Pose Graph Attention Network for Malnourishment Screening
by: Khan, Misaal, et al.
Published: (2025) -
D3: Data Diversity Design for Systematic Generalization in Visual Question Answering
by: Rahimi, Amir, et al.
Published: (2023)