A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zhenyang, Zheng, Sixiao, Chen, Siyu, Zhao, Cairong, Liang, Longfei, Xue, Xiangyang, Fu, Yanwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning
by: Liu, Zhenyang, et al.
Published: (2025)
by: Liu, Zhenyang, et al.
Published: (2025)
Spatial-Temporal Aware Visuomotor Diffusion Policy Learning
by: Liu, Zhenyang, et al.
Published: (2025)
by: Liu, Zhenyang, et al.
Published: (2025)
TriVLA: A Triple-System-Based Unified Vision-Language-Action Model with Episodic World Modeling for General Robot Control
by: Liu, Zhenyang, et al.
Published: (2025)
by: Liu, Zhenyang, et al.
Published: (2025)
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
by: Liu, Zhenyang, et al.
Published: (2026)
by: Liu, Zhenyang, et al.
Published: (2026)
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
by: Zheng, Sixiao, et al.
Published: (2024)
by: Zheng, Sixiao, et al.
Published: (2024)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
VADF: Vision-Adaptive Diffusion Policy Framework for Efficient Robotic Manipulation
by: Yu, Xinglei, et al.
Published: (2026)
by: Yu, Xinglei, et al.
Published: (2026)
OVGNet: A Unified Visual-Linguistic Framework for Open-Vocabulary Robotic Grasping
by: Meng, Li, et al.
Published: (2024)
by: Meng, Li, et al.
Published: (2024)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
by: Li, Danyang, et al.
Published: (2025)
by: Li, Danyang, et al.
Published: (2025)
Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance
by: Wang, Runze, et al.
Published: (2026)
by: Wang, Runze, et al.
Published: (2026)
OVAL-Prompt: Open-Vocabulary Affordance Localization for Robot Manipulation through LLM Affordance-Grounding
by: Tong, Edmond, et al.
Published: (2024)
by: Tong, Edmond, et al.
Published: (2024)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
by: Wang, Kuanning, et al.
Published: (2026)
by: Wang, Kuanning, et al.
Published: (2026)
Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image
by: Huang, Jingshun, et al.
Published: (2025)
by: Huang, Jingshun, et al.
Published: (2025)
Topo-Field: Topometric mapping with Brain-inspired Hierarchical Layout-Object-Position Fields
by: Hou, Jiawei, et al.
Published: (2024)
by: Hou, Jiawei, et al.
Published: (2024)
TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making
by: Li, Shanshan, et al.
Published: (2025)
by: Li, Shanshan, et al.
Published: (2025)
UniDiffGrasp: A Unified Framework Integrating VLM Reasoning and VLM-Guided Part Diffusion for Open-Vocabulary Constrained Grasping with Dual Arms
by: Guo, Xueyang, et al.
Published: (2025)
by: Guo, Xueyang, et al.
Published: (2025)
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
by: Hu, Xinggang, et al.
Published: (2026)
by: Hu, Xinggang, et al.
Published: (2026)
You Only Estimate Once: Unified, One-stage, Real-Time Category-level Articulated Object 6D Pose Estimation for Robotic Grasping
by: Huang, Jingshun, et al.
Published: (2025)
by: Huang, Jingshun, et al.
Published: (2025)
OFlow: Injecting Object-Aware Temporal Flow Matching for Robust Robotic Manipulation
by: Wang, Kuanning, et al.
Published: (2026)
by: Wang, Kuanning, et al.
Published: (2026)
OpenOcc: Open Vocabulary 3D Scene Reconstruction via Occupancy Representation
by: Jiang, Haochen, et al.
Published: (2024)
by: Jiang, Haochen, et al.
Published: (2024)
Intelligent Director: An Automatic Framework for Dynamic Visual Composition using ChatGPT
by: Zheng, Sixiao, et al.
Published: (2024)
by: Zheng, Sixiao, et al.
Published: (2024)
HiFi-CS: Towards Open Vocabulary Visual Grounding For Robotic Grasping Using Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2024)
by: Bhat, Vineet, et al.
Published: (2024)
SCOOP'D: Learning Mixed-Liquid-Solid Scooping via Sim2Real Generative Policy
by: Wang, Kuanning, et al.
Published: (2025)
by: Wang, Kuanning, et al.
Published: (2025)
Exploring Spatial Representation to Enhance LLM Reasoning in Aerial Vision-Language Navigation
by: Gao, Yunpeng, et al.
Published: (2024)
by: Gao, Yunpeng, et al.
Published: (2024)
Re$^2$MoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement
by: Zheng, Jiakun, et al.
Published: (2026)
by: Zheng, Jiakun, et al.
Published: (2026)
GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping
by: Ma, Teli, et al.
Published: (2024)
by: Ma, Teli, et al.
Published: (2024)
Learning Geometrically-Grounded 3D Visual Representations for View-Generalizable Robotic Manipulation
by: Zhang, Di, et al.
Published: (2026)
by: Zhang, Di, et al.
Published: (2026)
Language-Grounded Decoupled Action Representation for Robotic Manipulation
by: Weng, Wuding, et al.
Published: (2026)
by: Weng, Wuding, et al.
Published: (2026)
Grounded Vision-Language Navigation for UAVs with Open-Vocabulary Goal Understanding
by: Zhang, Yuhang, et al.
Published: (2025)
by: Zhang, Yuhang, et al.
Published: (2025)
OpenGraph: Open-Vocabulary Hierarchical 3D Graph Representation in Large-Scale Outdoor Environments
by: Deng, Yinan, et al.
Published: (2024)
by: Deng, Yinan, et al.
Published: (2024)
AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning
by: Kong, Yangzhe, et al.
Published: (2025)
by: Kong, Yangzhe, et al.
Published: (2025)
OV9D: Open-Vocabulary Category-Level 9D Object Pose and Size Estimation
by: Cai, Junhao, et al.
Published: (2024)
by: Cai, Junhao, et al.
Published: (2024)
OpenLex3D: A Tiered Evaluation Benchmark for Open-Vocabulary 3D Scene Representations
by: Kassab, Christina, et al.
Published: (2025)
by: Kassab, Christina, et al.
Published: (2025)
OpenIN: Open-Vocabulary Instance-Oriented Navigation in Dynamic Domestic Environments
by: Tang, Yujie, et al.
Published: (2025)
by: Tang, Yujie, et al.
Published: (2025)
osmAG-LLM: Zero-Shot Open-Vocabulary Object Navigation via Semantic Maps and Large Language Models Reasoning
by: Xie, Fujing, et al.
Published: (2025)
by: Xie, Fujing, et al.
Published: (2025)
DRIVE-Nav: Directional Reasoning, Inspection, and Verification for Efficient Open-Vocabulary Navigation
by: Gao, Maoguo, et al.
Published: (2026)
by: Gao, Maoguo, et al.
Published: (2026)
DIV-Nav: Open-Vocabulary Spatial Relationships for Multi-Object Navigation
by: Ortega-Peimbert, Jesús, et al.
Published: (2025)
by: Ortega-Peimbert, Jesús, et al.
Published: (2025)
AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
by: Tang, Yingbo, et al.
Published: (2025)
by: Tang, Yingbo, et al.
Published: (2025)
Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization
by: Yang, Jonathan, et al.
Published: (2025)
by: Yang, Jonathan, et al.
Published: (2025)
Similar Items
-
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning
by: Liu, Zhenyang, et al.
Published: (2025) -
Spatial-Temporal Aware Visuomotor Diffusion Policy Learning
by: Liu, Zhenyang, et al.
Published: (2025) -
TriVLA: A Triple-System-Based Unified Vision-Language-Action Model with Episodic World Modeling for General Robot Control
by: Liu, Zhenyang, et al.
Published: (2025) -
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
by: Liu, Zhenyang, et al.
Published: (2026) -
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
by: Zheng, Sixiao, et al.
Published: (2024)