Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kabir, Imran, Reza, Md Alimoor, Billah, Syed |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2025)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2025)
Identifying Crucial Objects in Blind and Low-Vision Individuals' Navigation
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
A Dataset for Crucial Object Recognition in Blind and Low-Vision Individuals' Navigation
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
von: Korekata, Ryosuke, et al.
Veröffentlicht: (2025)
von: Korekata, Ryosuke, et al.
Veröffentlicht: (2025)
SignRAG: A Retrieval-Augmented System for Scalable Zero-Shot Road Sign Recognition
von: Zhu, Minghao, et al.
Veröffentlicht: (2025)
von: Zhu, Minghao, et al.
Veröffentlicht: (2025)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
von: Chen, Yi, et al.
Veröffentlicht: (2023)
von: Chen, Yi, et al.
Veröffentlicht: (2023)
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
von: Man, Yunze, et al.
Veröffentlicht: (2024)
von: Man, Yunze, et al.
Veröffentlicht: (2024)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
von: Song, Chan Hee, et al.
Veröffentlicht: (2024)
von: Song, Chan Hee, et al.
Veröffentlicht: (2024)
Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
von: Wu, Yin, et al.
Veröffentlicht: (2025)
von: Wu, Yin, et al.
Veröffentlicht: (2025)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
von: Elhenawy, Mohammed, et al.
Veröffentlicht: (2025)
Scene-Graph ViT: End-to-End Open-Vocabulary Visual Relationship Detection
von: Salzmann, Tim, et al.
Veröffentlicht: (2024)
von: Salzmann, Tim, et al.
Veröffentlicht: (2024)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
von: Sun, Qi, et al.
Veröffentlicht: (2024)
von: Sun, Qi, et al.
Veröffentlicht: (2024)
RoadFormer: Duplex Transformer for RGB-Normal Semantic Road Scene Parsing
von: Li, Jiahang, et al.
Veröffentlicht: (2023)
von: Li, Jiahang, et al.
Veröffentlicht: (2023)
Acoustic Field Video for Multimodal Scene Understanding
von: Kim, Daehwa, et al.
Veröffentlicht: (2026)
von: Kim, Daehwa, et al.
Veröffentlicht: (2026)
RoboUniView: Visual-Language Model with Unified View Representation for Robotic Manipulation
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
von: Liu, Fanfan, et al.
Veröffentlicht: (2024)
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
von: Schäfer, Finn Rasmus, et al.
Veröffentlicht: (2026)
von: Schäfer, Finn Rasmus, et al.
Veröffentlicht: (2026)
Tracking Meets Large Multimodal Models for Driving Scenario Understanding
von: Ishaq, Ayesha, et al.
Veröffentlicht: (2025)
von: Ishaq, Ayesha, et al.
Veröffentlicht: (2025)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
PhyGrasp: Generalizing Robotic Grasping with Physics-informed Large Multimodal Models
von: Guo, Dingkun, et al.
Veröffentlicht: (2024)
von: Guo, Dingkun, et al.
Veröffentlicht: (2024)
A Superalignment Framework in Autonomous Driving with Large Language Models
von: Kong, Xiangrui, et al.
Veröffentlicht: (2024)
von: Kong, Xiangrui, et al.
Veröffentlicht: (2024)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
von: Nasiriany, Soroush, et al.
Veröffentlicht: (2024)
von: Nasiriany, Soroush, et al.
Veröffentlicht: (2024)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
von: Li, Jiaang, et al.
Veröffentlicht: (2025)
von: Li, Jiaang, et al.
Veröffentlicht: (2025)
VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages
von: Atuhurra, Jesse, et al.
Veröffentlicht: (2025)
von: Atuhurra, Jesse, et al.
Veröffentlicht: (2025)
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
von: Jia, Baoxiong, et al.
Veröffentlicht: (2024)
von: Jia, Baoxiong, et al.
Veröffentlicht: (2024)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
von: Wang, Qiuchen, et al.
Veröffentlicht: (2026)
von: Wang, Qiuchen, et al.
Veröffentlicht: (2026)
SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025)
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025)
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
von: Huang, Jiaxin, et al.
Veröffentlicht: (2025)
von: Huang, Jiaxin, et al.
Veröffentlicht: (2025)
ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments
von: Gong, Ziyang, et al.
Veröffentlicht: (2026)
von: Gong, Ziyang, et al.
Veröffentlicht: (2026)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
Enhancing Underwater Object Detection through Spatio-Temporal Analysis and Spatial Attention Networks
von: Karri, Sai Likhith, et al.
Veröffentlicht: (2025)
von: Karri, Sai Likhith, et al.
Veröffentlicht: (2025)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
von: Chia, Yew Ken, et al.
Veröffentlicht: (2024)
GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration
von: Wake, Naoki, et al.
Veröffentlicht: (2023)
von: Wake, Naoki, et al.
Veröffentlicht: (2023)
VINGS-Mono: Visual-Inertial Gaussian Splatting Monocular SLAM in Large Scenes
von: Wu, Ke, et al.
Veröffentlicht: (2025)
von: Wu, Ke, et al.
Veröffentlicht: (2025)
DM2RM: Dual-Mode Multimodal Ranking for Target Objects and Receptacles Based on Open-Vocabulary Instructions
von: Korekata, Ryosuke, et al.
Veröffentlicht: (2024)
von: Korekata, Ryosuke, et al.
Veröffentlicht: (2024)
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
von: Ma, Boyi, et al.
Veröffentlicht: (2025)
von: Ma, Boyi, et al.
Veröffentlicht: (2025)
Embodied Scene Understanding for Vision Language Models via MetaVQA
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2025) -
Identifying Crucial Objects in Blind and Low-Vision Individuals' Navigation
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024) -
A Dataset for Crucial Object Recognition in Blind and Low-Vision Individuals' Navigation
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024) -
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
von: Wang, Zhaowei, et al.
Veröffentlicht: (2024) -
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
von: Korekata, Ryosuke, et al.
Veröffentlicht: (2025)