RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tan, Xichen, Ye, Yunfan, Luo, Yuanjing, Wan, Qian, Liu, Fang, Cai, Zhiping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ALLVB: All-in-One Long Video Understanding Benchmark
von: Tan, Xichen, et al.
Veröffentlicht: (2025)
von: Tan, Xichen, et al.
Veröffentlicht: (2025)
HOCA-Bench: Beyond Semantic Perception to Predictive World Modeling via Hegelian Ontological-Causal Anomalies
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
von: Xue, Zhucun, et al.
Veröffentlicht: (2025)
Trans-Adapter: A Plug-and-Play Framework for Transparent Image Inpainting
von: Dai, Yuekun, et al.
Veröffentlicht: (2025)
von: Dai, Yuekun, et al.
Veröffentlicht: (2025)
Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2024)
TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
von: Cao, Zongsheng, et al.
Veröffentlicht: (2025)
von: Cao, Zongsheng, et al.
Veröffentlicht: (2025)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
HumanSAM: Classifying Human-centric Forgery Videos in Human Spatial, Appearance, and Motion Anomaly
von: Liu, Chang, et al.
Veröffentlicht: (2025)
von: Liu, Chang, et al.
Veröffentlicht: (2025)
StyleGallery: Training-free and Semantic-aware Personalized Style Transfer from Arbitrary Image References
von: He, Boyu, et al.
Veröffentlicht: (2026)
von: He, Boyu, et al.
Veröffentlicht: (2026)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
NVS-Adapter: Plug-and-Play Novel View Synthesis from a Single Image
von: Jeong, Yoonwoo, et al.
Veröffentlicht: (2023)
von: Jeong, Yoonwoo, et al.
Veröffentlicht: (2023)
Plug-and-Play Versatile Compressed Video Enhancement
von: Zeng, Huimin, et al.
Veröffentlicht: (2025)
von: Zeng, Huimin, et al.
Veröffentlicht: (2025)
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
von: Xie, Peijin, et al.
Veröffentlicht: (2025)
von: Xie, Peijin, et al.
Veröffentlicht: (2025)
SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
von: Chen, Jian, et al.
Veröffentlicht: (2024)
von: Chen, Jian, et al.
Veröffentlicht: (2024)
Protecting NeRFs' Copyright via Plug-And-Play Watermarking Base Model
von: Song, Qi, et al.
Veröffentlicht: (2024)
von: Song, Qi, et al.
Veröffentlicht: (2024)
DiffusionEdge: Diffusion Probabilistic Model for Crisp Edge Detection
von: Ye, Yunfan, et al.
Veröffentlicht: (2024)
von: Ye, Yunfan, et al.
Veröffentlicht: (2024)
ReMA: A Training-Free Plug-and-Play Mixing Augmentation for Video Behavior Recognition
von: Cui, Feng-Qi, et al.
Veröffentlicht: (2026)
von: Cui, Feng-Qi, et al.
Veröffentlicht: (2026)
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
von: Zhu, Lunjie, et al.
Veröffentlicht: (2026)
von: Zhu, Lunjie, et al.
Veröffentlicht: (2026)
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
von: Luo, Katie, et al.
Veröffentlicht: (2025)
von: Luo, Katie, et al.
Veröffentlicht: (2025)
Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
von: Xue, Bowen, et al.
Veröffentlicht: (2025)
von: Xue, Bowen, et al.
Veröffentlicht: (2025)
Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters
von: Chiu, Pin-Yen, et al.
Veröffentlicht: (2025)
von: Chiu, Pin-Yen, et al.
Veröffentlicht: (2025)
SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
von: Zeng, Nianbo, et al.
Veröffentlicht: (2025)
von: Zeng, Nianbo, et al.
Veröffentlicht: (2025)
Plug-and-Play Diffusion Distillation
von: Hsiao, Yi-Ting, et al.
Veröffentlicht: (2024)
von: Hsiao, Yi-Ting, et al.
Veröffentlicht: (2024)
Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning
von: Yang, Songyuan, et al.
Veröffentlicht: (2026)
von: Yang, Songyuan, et al.
Veröffentlicht: (2026)
iRAG: Advancing RAG for Videos with an Incremental Approach
von: Arefeen, Md Adnan, et al.
Veröffentlicht: (2024)
von: Arefeen, Md Adnan, et al.
Veröffentlicht: (2024)
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios
von: Yan, Peizheng, et al.
Veröffentlicht: (2026)
von: Yan, Peizheng, et al.
Veröffentlicht: (2026)
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
von: Li, Yinglu, et al.
Veröffentlicht: (2025)
von: Li, Yinglu, et al.
Veröffentlicht: (2025)
From Camera to World: A Plug-and-Play Module for Human Mesh Transformation
von: Ma, Changhai, et al.
Veröffentlicht: (2025)
von: Ma, Changhai, et al.
Veröffentlicht: (2025)
Plug-and-Play Logit Fusion for Heterogeneous Pathology Foundation Models
von: Huang, Gexin, et al.
Veröffentlicht: (2026)
von: Huang, Gexin, et al.
Veröffentlicht: (2026)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
von: Ren, Xubin, et al.
Veröffentlicht: (2025)
QualiRAG: Retrieval-Augmented Generation for Visual Quality Understanding
von: Cao, Linhan, et al.
Veröffentlicht: (2026)
von: Cao, Linhan, et al.
Veröffentlicht: (2026)
SpeedUpNet: A Plug-and-Play Adapter Network for Accelerating Text-to-Image Diffusion Models
von: Chai, Weilong, et al.
Veröffentlicht: (2023)
von: Chai, Weilong, et al.
Veröffentlicht: (2023)
Score-Based Turbo Message Passing for Plug-and-Play Compressive Imaging
von: Cai, Chang, et al.
Veröffentlicht: (2025)
von: Cai, Chang, et al.
Veröffentlicht: (2025)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
Text-to-Image Rectified Flow as Plug-and-Play Priors
von: Yang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Yang, Xiaofeng, et al.
Veröffentlicht: (2024)
Plug-and-Play Context Feature Reuse for Efficient Masked Generation
von: Liu, Xuejie, et al.
Veröffentlicht: (2025)
von: Liu, Xuejie, et al.
Veröffentlicht: (2025)
Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs
von: Liu, Jiazhen, et al.
Veröffentlicht: (2025)
von: Liu, Jiazhen, et al.
Veröffentlicht: (2025)
PnP-U3D: Plug-and-Play 3D Framework Bridging Autoregression and Diffusion for Unified Understanding and Generation
von: Chen, Yongwei, et al.
Veröffentlicht: (2026)
von: Chen, Yongwei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ALLVB: All-in-One Long Video Understanding Benchmark
von: Tan, Xichen, et al.
Veröffentlicht: (2025) -
HOCA-Bench: Beyond Semantic Perception to Predictive World Modeling via Hegelian Ontological-Causal Anomalies
von: Liu, Chang, et al.
Veröffentlicht: (2026) -
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
von: Xue, Zhucun, et al.
Veröffentlicht: (2025) -
Trans-Adapter: A Plug-and-Play Framework for Transparent Image Inpainting
von: Dai, Yuekun, et al.
Veröffentlicht: (2025) -
Generalizing Deepfake Video Detection with Plug-and-Play: Video-Level Blending and Spatiotemporal Adapter Tuning
von: Yan, Zhiyuan, et al.
Veröffentlicht: (2024)