E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xu, Zeyu, Zhang, Junkang, Wang, Qiang, Liu, Yi |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
VRAG-DFD: Verifiable Retrieval-Augmentation for MLLM-based Deepfake Detection
par: Han, Hui, et autres
Publié: (2026)
par: Han, Hui, et autres
Publié: (2026)
VRAG: Learning World Models for Interactive Video Generation
par: Chen, Taiye, et autres
Publié: (2025)
par: Chen, Taiye, et autres
Publié: (2025)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
par: Xue, Zhucun, et autres
Publié: (2025)
par: Xue, Zhucun, et autres
Publié: (2025)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
par: Shen, Xiaoqian, et autres
Publié: (2025)
par: Shen, Xiaoqian, et autres
Publié: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
par: Yuan, Huaying, et autres
Publié: (2025)
par: Yuan, Huaying, et autres
Publié: (2025)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
par: Peruzzo, Elia, et autres
Publié: (2025)
par: Peruzzo, Elia, et autres
Publié: (2025)
Long Video Understanding with Learnable Retrieval in Video-Language Models
par: Xu, Jiaqi, et autres
Publié: (2023)
par: Xu, Jiaqi, et autres
Publié: (2023)
ReMoMask: Retrieval-Augmented Masked Motion Generation
par: Li, Zhengdao, et autres
Publié: (2025)
par: Li, Zhengdao, et autres
Publié: (2025)
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
par: Ren, Weiming, et autres
Publié: (2024)
par: Ren, Weiming, et autres
Publié: (2024)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
par: Zhang, Yunzhu, et autres
Publié: (2025)
par: Zhang, Yunzhu, et autres
Publié: (2025)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
par: Chen, Yuxiao, et autres
Publié: (2026)
par: Chen, Yuxiao, et autres
Publié: (2026)
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
par: Ma, Wufei, et autres
Publié: (2024)
par: Ma, Wufei, et autres
Publié: (2024)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
par: Ren, Xubin, et autres
Publié: (2025)
par: Ren, Xubin, et autres
Publié: (2025)
PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing
par: Li, Peize, et autres
Publié: (2026)
par: Li, Peize, et autres
Publié: (2026)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
par: Jin, Hongbo, et autres
Publié: (2025)
par: Jin, Hongbo, et autres
Publié: (2025)
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
par: Chen, Jian, et autres
Publié: (2025)
par: Chen, Jian, et autres
Publié: (2025)
DrVideo: Document Retrieval Based Long Video Understanding
par: Ma, Ziyu, et autres
Publié: (2024)
par: Ma, Ziyu, et autres
Publié: (2024)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
par: Zhang, Haoji, et autres
Publié: (2025)
par: Zhang, Haoji, et autres
Publié: (2025)
Memory-Augmented Query Intent Understanding for Efficient Chat-based Image Retrieval
par: Chen, Xianke, et autres
Publié: (2026)
par: Chen, Xianke, et autres
Publié: (2026)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
par: Hu, Pengfei, et autres
Publié: (2025)
par: Hu, Pengfei, et autres
Publié: (2025)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
par: Li, Handong, et autres
Publié: (2026)
par: Li, Handong, et autres
Publié: (2026)
Retrieval-Augmented Egocentric Video Captioning
par: Xu, Jilan, et autres
Publié: (2024)
par: Xu, Jilan, et autres
Publié: (2024)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
par: Gao, Zhe, et autres
Publié: (2026)
par: Gao, Zhe, et autres
Publié: (2026)
QualiRAG: Retrieval-Augmented Generation for Visual Quality Understanding
par: Cao, Linhan, et autres
Publié: (2026)
par: Cao, Linhan, et autres
Publié: (2026)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
par: Zhu, Chenhui, et autres
Publié: (2025)
par: Zhu, Chenhui, et autres
Publié: (2025)
LVC: A Lightweight Compression Framework for Enhancing VLMs in Long Video Understanding
par: Wang, Ziyi, et autres
Publié: (2025)
par: Wang, Ziyi, et autres
Publié: (2025)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
par: Yu, Jiwen, et autres
Publié: (2025)
par: Yu, Jiwen, et autres
Publié: (2025)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
par: Huang, De-An, et autres
Publié: (2025)
par: Huang, De-An, et autres
Publié: (2025)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
par: Zeng, Nianbo, et autres
Publié: (2025)
par: Zeng, Nianbo, et autres
Publié: (2025)
Generative Frame Sampler for Long Video Understanding
par: Yao, Linli, et autres
Publié: (2025)
par: Yao, Linli, et autres
Publié: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
par: Qiu, Zongyang, et autres
Publié: (2025)
par: Qiu, Zongyang, et autres
Publié: (2025)
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
par: Wang, Qiuchen, et autres
Publié: (2025)
par: Wang, Qiuchen, et autres
Publié: (2025)
Motion Mamba: Efficient and Long Sequence Motion Generation
par: Zhang, Zeyu, et autres
Publié: (2024)
par: Zhang, Zeyu, et autres
Publié: (2024)
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
par: Xu, Weili, et autres
Publié: (2025)
par: Xu, Weili, et autres
Publié: (2025)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
par: Zheng, Xu, et autres
Publié: (2025)
par: Zheng, Xu, et autres
Publié: (2025)
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
par: Li, Yinglu, et autres
Publié: (2025)
par: Li, Yinglu, et autres
Publié: (2025)
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
par: Nie, Ming, et autres
Publié: (2025)
par: Nie, Ming, et autres
Publié: (2025)
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
par: Wu, Peiran, et autres
Publié: (2025)
par: Wu, Peiran, et autres
Publié: (2025)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
par: Jiang, Jindong, et autres
Publié: (2025)
par: Jiang, Jindong, et autres
Publié: (2025)
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
par: Li, Jungang, et autres
Publié: (2024)
par: Li, Jungang, et autres
Publié: (2024)
Documents similaires
-
VRAG-DFD: Verifiable Retrieval-Augmentation for MLLM-based Deepfake Detection
par: Han, Hui, et autres
Publié: (2026) -
VRAG: Learning World Models for Interactive Video Generation
par: Chen, Taiye, et autres
Publié: (2025) -
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
par: Xue, Zhucun, et autres
Publié: (2025) -
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
par: Shen, Xiaoqian, et autres
Publié: (2025) -
Memory-enhanced Retrieval Augmentation for Long Video Understanding
par: Yuan, Huaying, et autres
Publié: (2025)