D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zihan, Lee, Seungjun, Dai, Guangzhao, Lee, Gim Hee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
by: Lee, Seungjun, et al.
Published: (2026)
by: Lee, Seungjun, et al.
Published: (2026)
Segment Any 3D Object with Language
by: Lee, Seungjun, et al.
Published: (2024)
by: Lee, Seungjun, et al.
Published: (2024)
DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splatting
by: Lee, Seungjun, et al.
Published: (2025)
by: Lee, Seungjun, et al.
Published: (2025)
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation
by: Zhou, Hanyu, et al.
Published: (2026)
by: Zhou, Hanyu, et al.
Published: (2026)
Segment Any Events with Language
by: Lee, Seungjun, et al.
Published: (2026)
by: Lee, Seungjun, et al.
Published: (2026)
Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
RiemanLine: Riemannian Manifold Representation of 3D Lines for Factor Graph Optimization
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene Understanding
by: Zhou, Haoran, et al.
Published: (2025)
by: Zhou, Haoran, et al.
Published: (2025)
Zero-Shot 3D Visual Grounding from Vision-Language Models
by: Li, Rong, et al.
Published: (2025)
by: Li, Rong, et al.
Published: (2025)
SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning
by: Yang, Yuyuan, et al.
Published: (2026)
by: Yang, Yuyuan, et al.
Published: (2026)
Open-Set 3D Semantic Instance Maps for Vision Language Navigation -- O3D-SIM
by: Nanwani, Laksh, et al.
Published: (2024)
by: Nanwani, Laksh, et al.
Published: (2024)
VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
by: Halacheva, Anna-Maria, et al.
Published: (2025)
by: Halacheva, Anna-Maria, et al.
Published: (2025)
IAAO: Interactive Affordance Learning for Articulated Objects in 3D Environments
by: Zhang, Can, et al.
Published: (2025)
by: Zhang, Can, et al.
Published: (2025)
Learning 3D Persistent Embodied World Models
by: Zhou, Siyuan, et al.
Published: (2025)
by: Zhou, Siyuan, et al.
Published: (2025)
MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
by: Huang, Xun, et al.
Published: (2025)
by: Huang, Xun, et al.
Published: (2025)
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
by: Yang, Yuncong, et al.
Published: (2024)
by: Yang, Yuncong, et al.
Published: (2024)
Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
by: Sohn, Tin Stribor, et al.
Published: (2025)
by: Sohn, Tin Stribor, et al.
Published: (2025)
HandMCM: Multi-modal Point Cloud-based Correspondence State Space Model for 3D Hand Pose Estimation
by: Cheng, Wencan, et al.
Published: (2026)
by: Cheng, Wencan, et al.
Published: (2026)
EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence
by: Wang, Xinjie, et al.
Published: (2025)
by: Wang, Xinjie, et al.
Published: (2025)
econSG: Efficient and Multi-view Consistent Open-Vocabulary 3D Semantic Gaussians
by: Zhang, Can, et al.
Published: (2025)
by: Zhang, Can, et al.
Published: (2025)
NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning
by: Ju, Yuanchen, et al.
Published: (2025)
by: Ju, Yuanchen, et al.
Published: (2025)
RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action
by: Sun, Xiaoquan, et al.
Published: (2025)
by: Sun, Xiaoquan, et al.
Published: (2025)
DOGS: Distributed-Oriented Gaussian Splatting for Large-Scale 3D Reconstruction Via Gaussian Consensus
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
by: Zhang, Jiyao, et al.
Published: (2026)
by: Zhang, Jiyao, et al.
Published: (2026)
ChatSplat: 3D Conversational Gaussian Splatting
by: Chen, Hanlin, et al.
Published: (2024)
by: Chen, Hanlin, et al.
Published: (2024)
SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation
by: Zhang, Can, et al.
Published: (2026)
by: Zhang, Can, et al.
Published: (2026)
ZING-3D: Zero-shot Incremental 3D Scene Graphs via Vision-Language Models
by: Saxena, Pranav, et al.
Published: (2025)
by: Saxena, Pranav, et al.
Published: (2025)
Syn-to-Real Unsupervised Domain Adaptation for Indoor 3D Object Detection
by: Wang, Yunsong, et al.
Published: (2024)
by: Wang, Yunsong, et al.
Published: (2024)
Enhancing Generalizability of Representation Learning for Data-Efficient 3D Scene Understanding
by: Wang, Yunsong, et al.
Published: (2024)
by: Wang, Yunsong, et al.
Published: (2024)
LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
MotionScale: Reconstructing Appearance, Geometry, and Motion of Dynamic Scenes with Scalable 4D Gaussian Splatting
by: Zhou, Haoran, et al.
Published: (2026)
by: Zhou, Haoran, et al.
Published: (2026)
3D Reconstruction-Based Seed Counting of Sorghum Panicles for Agricultural Inspection
by: Freeman, Harry, et al.
Published: (2022)
by: Freeman, Harry, et al.
Published: (2022)
TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size
by: Lionar, Stefan, et al.
Published: (2026)
by: Lionar, Stefan, et al.
Published: (2026)
Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
by: Zhou, Hanyu, et al.
Published: (2025)
by: Zhou, Hanyu, et al.
Published: (2025)
3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
Similar Items
-
Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation
by: Wang, Zihan, et al.
Published: (2025) -
g3D-LF: Generalizable 3D-Language Feature Fields for Embodied Tasks
by: Wang, Zihan, et al.
Published: (2024) -
EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
by: Lee, Seungjun, et al.
Published: (2026) -
Segment Any 3D Object with Language
by: Lee, Seungjun, et al.
Published: (2024) -
DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splatting
by: Lee, Seungjun, et al.
Published: (2025)