A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
Fuente:
arXiv
Saved in:
| Main Authors: | Jahangard, Simindokht, Mohammadi, Mehrzad, Dhall, Abhinav, Rezatofighi, Hamid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025)
by: Jahangard, Simindokht, et al.
Published: (2025)
NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning
by: Cai, Zhixi, et al.
Published: (2025)
by: Cai, Zhixi, et al.
Published: (2025)
JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups
by: Jahangard, Simindokht, et al.
Published: (2024)
by: Jahangard, Simindokht, et al.
Published: (2024)
NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UAV Search Missions
by: Cai, Zhixi, et al.
Published: (2024)
by: Cai, Zhixi, et al.
Published: (2024)
HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning
by: Ke, Fucai, et al.
Published: (2024)
by: Ke, Fucai, et al.
Published: (2024)
VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning
by: Liang, Yichao, et al.
Published: (2024)
by: Liang, Yichao, et al.
Published: (2024)
JARVIS: A Neuro-Symbolic Commonsense Reasoning Framework for Conversational Embodied Agents
by: Zheng, Kaizhi, et al.
Published: (2022)
by: Zheng, Kaizhi, et al.
Published: (2022)
Imperative Learning: A Self-supervised Neuro-Symbolic Learning Framework for Robot Autonomy
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
JRDB-Pose3D: A Multi-person 3D Human Pose and Shape Estimation Dataset for Robotics
by: Biswas, Sandika, et al.
Published: (2026)
by: Biswas, Sandika, et al.
Published: (2026)
Multi-Modal World Model for Physical Robot Interactions: Simultaneous Visual and Tactile Predictions for Enhanced Accuracy
by: Mandil, Willow, et al.
Published: (2023)
by: Mandil, Willow, et al.
Published: (2023)
Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
by: Chia, Yew Ken, et al.
Published: (2024)
by: Chia, Yew Ken, et al.
Published: (2024)
Spatially Visual Perception for End-to-End Robotic Learning
by: Davies, Travis, et al.
Published: (2024)
by: Davies, Travis, et al.
Published: (2024)
SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
by: Qi, Zekun, et al.
Published: (2025)
by: Qi, Zekun, et al.
Published: (2025)
Improving Visual Perception of a Social Robot for Controlled and In-the-wild Human-robot Interaction
by: Zhong, Wangjie, et al.
Published: (2024)
by: Zhong, Wangjie, et al.
Published: (2024)
Neuro-Symbolic Concepts
by: Mao, Jiayuan, et al.
Published: (2025)
by: Mao, Jiayuan, et al.
Published: (2025)
Semantic-Drive: Democratizing Long-Tail Data Curation via Open-Vocabulary Grounding and Neuro-Symbolic VLM Consensus
by: Guillen-Perez, Antonio
Published: (2025)
by: Guillen-Perez, Antonio
Published: (2025)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations
by: Lyu, Ruiyuan, et al.
Published: (2024)
by: Lyu, Ruiyuan, et al.
Published: (2024)
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
by: Zahid, Azizul, et al.
Published: (2025)
by: Zahid, Azizul, et al.
Published: (2025)
P3-PO: Prescriptive Point Priors for Visuo-Spatial Generalization of Robot Policies
by: Levy, Mara, et al.
Published: (2024)
by: Levy, Mara, et al.
Published: (2024)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
SpatialAnt: Autonomous Zero-Shot Robot Navigation via Active Scene Reconstruction and Visual Anticipation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
by: Zantout, Nader, et al.
Published: (2025)
by: Zantout, Nader, et al.
Published: (2025)
Robotic Visual Instruction
by: Li, Yanbang, et al.
Published: (2025)
by: Li, Yanbang, et al.
Published: (2025)
Physically Grounded Vision-Language Models for Robotic Manipulation
by: Gao, Jensen, et al.
Published: (2023)
by: Gao, Jensen, et al.
Published: (2023)
iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
by: Wang, Kohou, et al.
Published: (2025)
by: Wang, Kohou, et al.
Published: (2025)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
by: Fang, Jiading
Published: (2025)
by: Fang, Jiading
Published: (2025)
MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning
by: Cai, Zhixi, et al.
Published: (2026)
by: Cai, Zhixi, et al.
Published: (2026)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
by: Chen, Shizhe, et al.
Published: (2025)
by: Chen, Shizhe, et al.
Published: (2025)
Tag Map: A Text-Based Map for Spatial Reasoning and Navigation with Large Language Models
by: Zhang, Mike, et al.
Published: (2024)
by: Zhang, Mike, et al.
Published: (2024)
Distracted Robot: How Visual Clutter Undermine Robotic Manipulation
by: Rasouli, Amir, et al.
Published: (2025)
by: Rasouli, Amir, et al.
Published: (2025)
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
by: Wang, Boyang, et al.
Published: (2026)
by: Wang, Boyang, et al.
Published: (2026)
Residual Cross-Modal Fusion Networks for Audio-Visual Navigation
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
SEM: Enhancing Spatial Understanding for Robust Robot Manipulation
by: Lin, Xuewu, et al.
Published: (2025)
by: Lin, Xuewu, et al.
Published: (2025)
Visual Foresight for Robotic Stow: A Diffusion-Based World Model from Sparse Snapshots
by: Zhang, Lijun, et al.
Published: (2026)
by: Zhang, Lijun, et al.
Published: (2026)
Multi-Scale Neighborhood Occupancy Masked Autoencoder for Self-Supervised Learning in LiDAR Point Clouds
by: Abdelsamad, Mohamed, et al.
Published: (2025)
by: Abdelsamad, Mohamed, et al.
Published: (2025)
Dynamic Robot-Assisted Surgery with Hierarchical Class-Incremental Semantic Segmentation
by: Hindel, Julia, et al.
Published: (2025)
by: Hindel, Julia, et al.
Published: (2025)
PhotoAgent: A Robotic Photographer with Spatial and Aesthetic Understanding
by: Che, Lirong, et al.
Published: (2026)
by: Che, Lirong, et al.
Published: (2026)
Visual IRL for Human-Like Robotic Manipulation
by: Asali, Ehsan, et al.
Published: (2024)
by: Asali, Ehsan, et al.
Published: (2024)
Similar Items
-
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025) -
NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning
by: Cai, Zhixi, et al.
Published: (2025) -
JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups
by: Jahangard, Simindokht, et al.
Published: (2024) -
NEUSIS: A Compositional Neuro-Symbolic Framework for Autonomous Perception, Reasoning, and Planning in Complex UAV Search Missions
by: Cai, Zhixi, et al.
Published: (2024) -
HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning
by: Ke, Fucai, et al.
Published: (2024)