D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zihan, Lee, Seungjun, Dai, Guangzhao, Lee, Gim Hee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909962153754624
author Wang, Zihan
Lee, Seungjun
Dai, Guangzhao
Lee, Gim Hee
author_facet Wang, Zihan
Lee, Seungjun
Dai, Guangzhao
Lee, Gim Hee
contents Embodied agents face a critical dilemma that end-to-end models lack interpretability and explicit 3D reasoning, while modular systems ignore cross-component interdependencies and synergies. To bridge this gap, we propose the Dynamic 3D Vision-Language-Planning Model (D3D-VLP). Our model introduces two key innovations: 1) A Dynamic 3D Chain-of-Thought (3D CoT) that unifies planning, grounding, navigation, and question answering within a single 3D-VLM and CoT pipeline; 2) A Synergistic Learning from Fragmented Supervision (SLFS) strategy, which uses a masked autoregressive loss to learn from massive and partially-annotated hybrid data. This allows different CoT components to mutually reinforce and implicitly supervise each other. To this end, we construct a large-scale dataset with 10M hybrid samples from 5K real scans and 20K synthetic scenes that are compatible with online learning methods such as RL and DAgger. Our D3D-VLP achieves state-of-the-art results on multiple benchmarks, including Vision-and-Language Navigation (R2R-CE, REVERIE-CE, NavRAG-CE), Object-goal Navigation (HM3D-OVON), and Task-oriented Sequential Grounding and Navigation (SG3D). Real-world mobile manipulation experiments further validate the effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12622
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
Wang, Zihan
Lee, Seungjun
Dai, Guangzhao
Lee, Gim Hee
Computer Vision and Pattern Recognition
Robotics
Embodied agents face a critical dilemma that end-to-end models lack interpretability and explicit 3D reasoning, while modular systems ignore cross-component interdependencies and synergies. To bridge this gap, we propose the Dynamic 3D Vision-Language-Planning Model (D3D-VLP). Our model introduces two key innovations: 1) A Dynamic 3D Chain-of-Thought (3D CoT) that unifies planning, grounding, navigation, and question answering within a single 3D-VLM and CoT pipeline; 2) A Synergistic Learning from Fragmented Supervision (SLFS) strategy, which uses a masked autoregressive loss to learn from massive and partially-annotated hybrid data. This allows different CoT components to mutually reinforce and implicitly supervise each other. To this end, we construct a large-scale dataset with 10M hybrid samples from 5K real scans and 20K synthetic scenes that are compatible with online learning methods such as RL and DAgger. Our D3D-VLP achieves state-of-the-art results on multiple benchmarks, including Vision-and-Language Navigation (R2R-CE, REVERIE-CE, NavRAG-CE), Object-goal Navigation (HM3D-OVON), and Task-oriented Sequential Grounding and Navigation (SG3D). Real-world mobile manipulation experiments further validate the effectiveness.
title D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2512.12622