SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Zeng, Nianbo, Hou, Haowen, Yu, Fei Richard, Shi, Si, He, Ying Tiffany |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SSEditor: Controllable Mask-to-Scene Generation with Diffusion Model
di: Zheng, Haowen, et al.
Pubblicazione: (2024)
di: Zheng, Haowen, et al.
Pubblicazione: (2024)
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
di: Chang, Yue, et al.
Pubblicazione: (2026)
di: Chang, Yue, et al.
Pubblicazione: (2026)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
di: Luo, Yongdong, et al.
Pubblicazione: (2024)
di: Luo, Yongdong, et al.
Pubblicazione: (2024)
Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model
di: Zhang, Shengjun, et al.
Pubblicazione: (2025)
di: Zhang, Shengjun, et al.
Pubblicazione: (2025)
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
di: Wang, Shuai, et al.
Pubblicazione: (2025)
di: Wang, Shuai, et al.
Pubblicazione: (2025)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
di: Ren, Xubin, et al.
Pubblicazione: (2025)
di: Ren, Xubin, et al.
Pubblicazione: (2025)
Revisit Human-Scene Interaction via Space Occupancy
di: Liu, Xinpeng, et al.
Pubblicazione: (2023)
di: Liu, Xinpeng, et al.
Pubblicazione: (2023)
Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning
di: Yu, Fei, et al.
Pubblicazione: (2025)
di: Yu, Fei, et al.
Pubblicazione: (2025)
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios
di: Yan, Peizheng, et al.
Pubblicazione: (2026)
di: Yan, Peizheng, et al.
Pubblicazione: (2026)
SceneX: Procedural Controllable Large-scale Scene Generation
di: Zhou, Mengqi, et al.
Pubblicazione: (2024)
di: Zhou, Mengqi, et al.
Pubblicazione: (2024)
Lifting Unlabeled Internet-level Data for 3D Scene Understanding
di: Chen, Yixin, et al.
Pubblicazione: (2026)
di: Chen, Yixin, et al.
Pubblicazione: (2026)
Response-G1: Explicit Scene Graph Modeling for Proactive Streaming Video Understanding
di: Ma, Ke, et al.
Pubblicazione: (2026)
di: Ma, Ke, et al.
Pubblicazione: (2026)
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
di: Yang, Xuyi, et al.
Pubblicazione: (2025)
di: Yang, Xuyi, et al.
Pubblicazione: (2025)
MetaFind: Scene-Aware 3D Asset Retrieval for Coherent Metaverse Scene Generation
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
di: Pan, Zhenyu, et al.
Pubblicazione: (2025)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
di: Liu, Hanqing, et al.
Pubblicazione: (2026)
di: Liu, Hanqing, et al.
Pubblicazione: (2026)
Mesh RAG: Retrieval Augmentation for Autoregressive Mesh Generation
di: Sun, Xiatao, et al.
Pubblicazione: (2025)
di: Sun, Xiatao, et al.
Pubblicazione: (2025)
Attention Mechanism based Cognition-level Scene Understanding
di: Tang, Xuejiao, et al.
Pubblicazione: (2022)
di: Tang, Xuejiao, et al.
Pubblicazione: (2022)
FastV-RAG: Towards Fast and Fine-Grained Video QA with Retrieval-Augmented Generation
di: Li, Gen, et al.
Pubblicazione: (2026)
di: Li, Gen, et al.
Pubblicazione: (2026)
Evaluating Compositional Scene Understanding in Multimodal Generative Models
di: Fu, Shuhao, et al.
Pubblicazione: (2025)
di: Fu, Shuhao, et al.
Pubblicazione: (2025)
HexPlane Representation for 3D Semantic Scene Understanding
di: Chen, Zeren, et al.
Pubblicazione: (2025)
di: Chen, Zeren, et al.
Pubblicazione: (2025)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
di: Liang, Yujia, et al.
Pubblicazione: (2025)
di: Liang, Yujia, et al.
Pubblicazione: (2025)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
di: Li, Haoyuan, et al.
Pubblicazione: (2025)
di: Li, Haoyuan, et al.
Pubblicazione: (2025)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
di: Zeng, Gangyan, et al.
Pubblicazione: (2024)
di: Zeng, Gangyan, et al.
Pubblicazione: (2024)
SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation
di: Wang, Wenjia, et al.
Pubblicazione: (2024)
di: Wang, Wenjia, et al.
Pubblicazione: (2024)
MTMamba: Enhancing Multi-Task Dense Scene Understanding by Mamba-Based Decoders
di: Lin, Baijiong, et al.
Pubblicazione: (2024)
di: Lin, Baijiong, et al.
Pubblicazione: (2024)
VideoRAG: Retrieval-Augmented Generation over Video Corpus
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving
di: Zhao, Rui, et al.
Pubblicazione: (2026)
di: Zhao, Rui, et al.
Pubblicazione: (2026)
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
di: Li, Yinglu, et al.
Pubblicazione: (2025)
di: Li, Yinglu, et al.
Pubblicazione: (2025)
ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
di: Liu, Fangfu, et al.
Pubblicazione: (2024)
di: Liu, Fangfu, et al.
Pubblicazione: (2024)
RAG-HAR: Retrieval Augmented Generation-based Human Activity Recognition
di: Sivaroopan, Nirhoshan, et al.
Pubblicazione: (2025)
di: Sivaroopan, Nirhoshan, et al.
Pubblicazione: (2025)
Towards Holistic Surgical Scene Understanding
di: Valderrama, Natalia, et al.
Pubblicazione: (2022)
di: Valderrama, Natalia, et al.
Pubblicazione: (2022)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
di: Huang, Jincai, et al.
Pubblicazione: (2026)
di: Huang, Jincai, et al.
Pubblicazione: (2026)
MV-RAG: Retrieval Augmented Multiview Diffusion
di: Dayani, Yosef, et al.
Pubblicazione: (2025)
di: Dayani, Yosef, et al.
Pubblicazione: (2025)
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes
di: Rivera, Antonio Carlos, et al.
Pubblicazione: (2024)
di: Rivera, Antonio Carlos, et al.
Pubblicazione: (2024)
PaintScene4D: Consistent 4D Scene Generation from Text Prompts
di: Gupta, Vinayak, et al.
Pubblicazione: (2024)
di: Gupta, Vinayak, et al.
Pubblicazione: (2024)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
di: Wang, Xingrui, et al.
Pubblicazione: (2024)
di: Wang, Xingrui, et al.
Pubblicazione: (2024)
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
di: Xie, Yifan, et al.
Pubblicazione: (2025)
di: Xie, Yifan, et al.
Pubblicazione: (2025)
QualiRAG: Retrieval-Augmented Generation for Visual Quality Understanding
di: Cao, Linhan, et al.
Pubblicazione: (2026)
di: Cao, Linhan, et al.
Pubblicazione: (2026)
BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving
di: Tang, Tao, et al.
Pubblicazione: (2024)
di: Tang, Tao, et al.
Pubblicazione: (2024)
DOCTR: Disentangled Object-Centric Transformer for Point Scene Understanding
di: Yu, Xiaoxuan, et al.
Pubblicazione: (2024)
di: Yu, Xiaoxuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SSEditor: Controllable Mask-to-Scene Generation with Diffusion Model
di: Zheng, Haowen, et al.
Pubblicazione: (2024) -
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
di: Chang, Yue, et al.
Pubblicazione: (2026) -
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
di: Luo, Yongdong, et al.
Pubblicazione: (2024) -
Scene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion Model
di: Zhang, Shengjun, et al.
Pubblicazione: (2025) -
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
di: Wang, Shuai, et al.
Pubblicazione: (2025)