Saved in:
| Main Authors: | Zantout, Nader, Zhang, Haochen, Kachana, Pujith, Qiu, Jinkai, Chen, Guofei, Zhang, Ji, Wang, Wenshan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.18684 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes
by: Zhang, Haochen, et al.
Published: (2025)
by: Zhang, Haochen, et al.
Published: (2025)
VLA-3D: A Dataset for 3D Semantic Scene Understanding and Navigation
by: Zhang, Haochen, et al.
Published: (2024)
by: Zhang, Haochen, et al.
Published: (2024)
Interactive-FAR:Interactive, Fast and Adaptable Routing for Navigation Among Movable Obstacles in Complex Unknown Environments
by: He, Botao, et al.
Published: (2024)
by: He, Botao, et al.
Published: (2024)
ApexNav: An Adaptive Exploration Strategy for Zero-Shot Object Navigation with Target-centric Semantic Fusion
by: Zhang, Mingjie, et al.
Published: (2025)
by: Zhang, Mingjie, et al.
Published: (2025)
SysNav: Multi-Level Systematic Cooperation Enables Real-World, Cross-Embodiment Object Navigation
by: Zhu, Haokun, et al.
Published: (2026)
by: Zhu, Haokun, et al.
Published: (2026)
BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object Navigation
by: Zhou, Zibo, et al.
Published: (2025)
by: Zhou, Zibo, et al.
Published: (2025)
LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
by: Saxena, Pranav, et al.
Published: (2025)
by: Saxena, Pranav, et al.
Published: (2025)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction
by: Li, Samuel, et al.
Published: (2025)
by: Li, Samuel, et al.
Published: (2025)
STRIVE: Structured Representation Integrating VLM Reasoning for Efficient Object Navigation
by: Zhu, Haokun, et al.
Published: (2025)
by: Zhu, Haokun, et al.
Published: (2025)
GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning
by: Lu, Yiren, et al.
Published: (2026)
by: Lu, Yiren, et al.
Published: (2026)
Object-centric Reconstruction and Tracking of Dynamic Unknown Objects using 3D Gaussian Splatting
by: Barad, Kuldeep R, et al.
Published: (2024)
by: Barad, Kuldeep R, et al.
Published: (2024)
Zero-Shot 3D Visual Grounding from Vision-Language Models
by: Li, Rong, et al.
Published: (2025)
by: Li, Rong, et al.
Published: (2025)
Air-FAR: Fast and Adaptable Routing for Aerial Navigation in Large-scale Complex Unknown Environments
by: He, Botao, et al.
Published: (2024)
by: He, Botao, et al.
Published: (2024)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
by: Tian, Tongxuan, et al.
Published: (2025)
by: Tian, Tongxuan, et al.
Published: (2025)
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
by: Sun, Xuefei, et al.
Published: (2026)
by: Sun, Xuefei, et al.
Published: (2026)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
Multi-Floor Zero-Shot Object Navigation Policy
by: Zhang, Lingfeng, et al.
Published: (2024)
by: Zhang, Lingfeng, et al.
Published: (2024)
TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation
by: Patel, Manthan, et al.
Published: (2025)
by: Patel, Manthan, et al.
Published: (2025)
MinkSORT: A 3D deep feature extractor using sparse convolutions to improve 3D multi-object tracking in greenhouse tomato plants
by: Rapado-Rincon, David, et al.
Published: (2023)
by: Rapado-Rincon, David, et al.
Published: (2023)
NavDreamer: Video Models as Zero-Shot 3D Navigators
by: Huang, Xijie, et al.
Published: (2026)
by: Huang, Xijie, et al.
Published: (2026)
AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models
by: Huynh, Cuong, et al.
Published: (2026)
by: Huynh, Cuong, et al.
Published: (2026)
D$^3$Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement
by: Wang, Yixuan, et al.
Published: (2023)
by: Wang, Yixuan, et al.
Published: (2023)
TriHelper: Zero-Shot Object Navigation with Dynamic Assistance
by: Zhang, Lingfeng, et al.
Published: (2024)
by: Zhang, Lingfeng, et al.
Published: (2024)
Informative Object-centric Next Best View for Object-aware 3D Gaussian Splatting in Cluttered Scenes
by: Jeong, Seunghoon, et al.
Published: (2026)
by: Jeong, Seunghoon, et al.
Published: (2026)
Dream2Real: Zero-Shot 3D Object Rearrangement with Vision-Language Models
by: Kapelyukh, Ivan, et al.
Published: (2023)
by: Kapelyukh, Ivan, et al.
Published: (2023)
Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
by: Yu, Houjian, et al.
Published: (2025)
by: Yu, Houjian, et al.
Published: (2025)
Zero-Shot Robotic Manipulation via 3D Gaussian Splatting-Enhanced Multimodal Retrieval-Augmented Generation
by: Xie, Zilong, et al.
Published: (2026)
by: Xie, Zilong, et al.
Published: (2026)
Planning and Reasoning with 3D Deformable Objects for Hierarchical Text-to-3D Robotic Shaping
by: Bartsch, Alison, et al.
Published: (2024)
by: Bartsch, Alison, et al.
Published: (2024)
SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
by: Huang, Jiaxin, et al.
Published: (2025)
by: Huang, Jiaxin, et al.
Published: (2025)
Learning Surgical Robotic Manipulation with 3D Spatial Priors
by: Sheng, Yu, et al.
Published: (2026)
by: Sheng, Yu, et al.
Published: (2026)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Fly0: Decoupling Semantic Grounding from Geometric Planning for Zero-Shot Aerial Navigation
by: Xu, Zhenxing, et al.
Published: (2026)
by: Xu, Zhenxing, et al.
Published: (2026)
EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene Supervision
by: Chen, Jiahao, et al.
Published: (2026)
by: Chen, Jiahao, et al.
Published: (2026)
SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
by: Qi, Zekun, et al.
Published: (2025)
by: Qi, Zekun, et al.
Published: (2025)
osmAG-LLM: Zero-Shot Open-Vocabulary Object Navigation via Semantic Maps and Large Language Models Reasoning
by: Xie, Fujing, et al.
Published: (2025)
by: Xie, Fujing, et al.
Published: (2025)
DreamGrasp: Zero-Shot 3D Multi-Object Reconstruction from Partial-View Images for Robotic Manipulation
by: Kim, Young Hun, et al.
Published: (2025)
by: Kim, Young Hun, et al.
Published: (2025)
REFLEX: Metacognitive Reasoning for Reflective Zero-Shot Robotic Planning with Large Language Models
by: Lin, Wenjie, et al.
Published: (2025)
by: Lin, Wenjie, et al.
Published: (2025)
Object-centric 3D Motion Field for Robot Learning from Human Videos
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
Think, Remember, Navigate: Zero-Shot Object-Goal Navigation with VLM-Powered Reasoning
by: Habibpour, Mobin, et al.
Published: (2025)
by: Habibpour, Mobin, et al.
Published: (2025)
Similar Items
-
IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes
by: Zhang, Haochen, et al.
Published: (2025) -
VLA-3D: A Dataset for 3D Semantic Scene Understanding and Navigation
by: Zhang, Haochen, et al.
Published: (2024) -
Interactive-FAR:Interactive, Fast and Adaptable Routing for Navigation Among Movable Obstacles in Complex Unknown Environments
by: He, Botao, et al.
Published: (2024) -
ApexNav: An Adaptive Exploration Strategy for Zero-Shot Object Navigation with Target-centric Semantic Fusion
by: Zhang, Mingjie, et al.
Published: (2025) -
SysNav: Multi-Level Systematic Cooperation Enables Real-World, Cross-Embodiment Object Navigation
by: Zhu, Haokun, et al.
Published: (2026)