NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Zihan, Zhu, Yaohui, Lee, Gim Hee, Fan, Yachun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks
di: Wang, Zihan, et al.
Pubblicazione: (2024)
di: Wang, Zihan, et al.
Pubblicazione: (2024)
NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
di: Yang, Haolin, et al.
Pubblicazione: (2025)
di: Yang, Haolin, et al.
Pubblicazione: (2025)
EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning
di: Lin, Bingqian, et al.
Pubblicazione: (2025)
di: Lin, Bingqian, et al.
Pubblicazione: (2025)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
di: Wang, Qiuchen, et al.
Pubblicazione: (2026)
di: Wang, Qiuchen, et al.
Pubblicazione: (2026)
NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
di: He, Zongtao, et al.
Pubblicazione: (2025)
di: He, Zongtao, et al.
Pubblicazione: (2025)
Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
di: Lee, Seungjun, et al.
Pubblicazione: (2026)
di: Lee, Seungjun, et al.
Pubblicazione: (2026)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
di: Loo, Gowen, et al.
Pubblicazione: (2025)
di: Loo, Gowen, et al.
Pubblicazione: (2025)
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
di: Long, Yuxing, et al.
Pubblicazione: (2024)
di: Long, Yuxing, et al.
Pubblicazione: (2024)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
di: Gao, Bingjie, et al.
Pubblicazione: (2025)
di: Gao, Bingjie, et al.
Pubblicazione: (2025)
LaB-RAG: Label Boosted Retrieval Augmented Generation for Radiology Report Generation
di: Song, Steven, et al.
Pubblicazione: (2024)
di: Song, Steven, et al.
Pubblicazione: (2024)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
di: Yu, Zhuoyuan, et al.
Pubblicazione: (2025)
di: Yu, Zhuoyuan, et al.
Pubblicazione: (2025)
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
AIP: Subverting Retrieval-Augmented Generation via Adversarial Instructional Prompt
di: Chaturvedi, Saket S., et al.
Pubblicazione: (2025)
di: Chaturvedi, Saket S., et al.
Pubblicazione: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
di: Wu, Yin, et al.
Pubblicazione: (2025)
di: Wu, Yin, et al.
Pubblicazione: (2025)
Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
di: Zhou, Hanyu, et al.
Pubblicazione: (2025)
di: Zhou, Hanyu, et al.
Pubblicazione: (2025)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
di: Korekata, Ryosuke, et al.
Pubblicazione: (2025)
di: Korekata, Ryosuke, et al.
Pubblicazione: (2025)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
di: Sun, Yubo, et al.
Pubblicazione: (2025)
di: Sun, Yubo, et al.
Pubblicazione: (2025)
Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation
di: Zhu, Mengdan, et al.
Pubblicazione: (2025)
di: Zhu, Mengdan, et al.
Pubblicazione: (2025)
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation
di: Hu, Chan-Wei, et al.
Pubblicazione: (2025)
di: Hu, Chan-Wei, et al.
Pubblicazione: (2025)
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
di: Wang, Wenbin, et al.
Pubblicazione: (2025)
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
di: Zhou, Hanyu, et al.
Pubblicazione: (2026)
di: Zhou, Hanyu, et al.
Pubblicazione: (2026)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
di: Tanaka, Ryota, et al.
Pubblicazione: (2025)
di: Tanaka, Ryota, et al.
Pubblicazione: (2025)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
di: Hsiao, Chi-Hsiang, et al.
Pubblicazione: (2025)
di: Hsiao, Chi-Hsiang, et al.
Pubblicazione: (2025)
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
di: Lin, Bingqian, et al.
Pubblicazione: (2024)
di: Lin, Bingqian, et al.
Pubblicazione: (2024)
Nav-R1: Reasoning and Navigation in Embodied Scenes
di: Liu, Qingxiang, et al.
Pubblicazione: (2025)
di: Liu, Qingxiang, et al.
Pubblicazione: (2025)
LangNav: Language as a Perceptual Representation for Navigation
di: Pan, Bowen, et al.
Pubblicazione: (2023)
di: Pan, Bowen, et al.
Pubblicazione: (2023)
Unified Geometry and Color Compression Framework for Point Clouds via Generative Diffusion Priors
di: Huang, Tianxin, et al.
Pubblicazione: (2025)
di: Huang, Tianxin, et al.
Pubblicazione: (2025)
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
di: Zhao, Beidi, et al.
Pubblicazione: (2026)
di: Zhao, Beidi, et al.
Pubblicazione: (2026)
VideoRAG: Retrieval-Augmented Generation over Video Corpus
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
OctoNav: Towards Generalist Embodied Navigation
di: Gao, Chen, et al.
Pubblicazione: (2025)
di: Gao, Chen, et al.
Pubblicazione: (2025)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
di: Sarwar, Nobin
Pubblicazione: (2025)
di: Sarwar, Nobin
Pubblicazione: (2025)
NavBench: Probing Multimodal Large Language Models for Embodied Navigation
di: Qiao, Yanyuan, et al.
Pubblicazione: (2025)
di: Qiao, Yanyuan, et al.
Pubblicazione: (2025)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
di: Hao, Haihong, et al.
Pubblicazione: (2025)
di: Hao, Haihong, et al.
Pubblicazione: (2025)
Semantic Map-based Generation of Navigation Instructions
di: Li, Chengzu, et al.
Pubblicazione: (2024)
di: Li, Chengzu, et al.
Pubblicazione: (2024)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
di: Han, Donghoon, et al.
Pubblicazione: (2024)
di: Han, Donghoon, et al.
Pubblicazione: (2024)
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
Segment Any Events with Language
di: Lee, Seungjun, et al.
Pubblicazione: (2026)
di: Lee, Seungjun, et al.
Pubblicazione: (2026)
CLAIR: CLIP-Aided Weakly Supervised Zero-Shot Cross-Domain Image Retrieval
di: Tan, Chor Boon, et al.
Pubblicazione: (2025)
di: Tan, Chor Boon, et al.
Pubblicazione: (2025)
Documenti analoghi
-
D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
di: Wang, Zihan, et al.
Pubblicazione: (2025) -
g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks
di: Wang, Zihan, et al.
Pubblicazione: (2024) -
NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
di: Yang, Haolin, et al.
Pubblicazione: (2025) -
EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning
di: Lin, Bingqian, et al.
Pubblicazione: (2025) -
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
di: Wang, Qiuchen, et al.
Pubblicazione: (2026)