Grounding Vision and Language to 3D Masks for Long-Horizon Box Rearrangement

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Malik, Ashish, Lowe, Caleb, Shrestha, Aayam, Lee, Stefan, Li, Fuxin, Fern, Alan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908912057319424
author Malik, Ashish
Lowe, Caleb
Shrestha, Aayam
Lee, Stefan
Li, Fuxin
Fern, Alan
author_facet Malik, Ashish
Lowe, Caleb
Shrestha, Aayam
Lee, Stefan
Li, Fuxin
Fern, Alan
contents We study long-horizon planning in 3D environments from under-specified natural-language goals using only visual observations, focusing on multi-step 3D box rearrangement tasks. Existing approaches typically rely on symbolic planners with brittle relational grounding of states and goals, or on direct action-sequence generation from 2D vision-language models (VLMs). Both approaches struggle with reasoning over many objects, rich 3D geometry, and implicit semantic constraints. Recent advances in 3D VLMs demonstrate strong grounding of natural-language referents to 3D segmentation masks, suggesting the potential for more general planning capabilities. We extend existing 3D grounding models and propose Reactive Action Mask Planner (RAMP-3D), which formulates long-horizon planning as sequential reactive prediction of paired 3D masks: a "which-object" mask indicating what to pick and a "which-target-region" mask specifying where to place it. The resulting system processes RGB-D observations and natural-language task specifications to reactively generate multi-step pick-and-place actions for 3D box rearrangement. We conduct experiments across 11 task variants in warehouse-style environments with 1-30 boxes and diverse natural-language constraints. RAMP-3D achieves 79.5% success rate on long-horizon rearrangement tasks and significantly outperforms 2D VLM-based baselines, establishing mask-based reactive policies as a promising alternative to symbolic pipelines for long-horizon planning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23676
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Grounding Vision and Language to 3D Masks for Long-Horizon Box Rearrangement
Malik, Ashish
Lowe, Caleb
Shrestha, Aayam
Lee, Stefan
Li, Fuxin
Fern, Alan
Artificial Intelligence
Robotics
We study long-horizon planning in 3D environments from under-specified natural-language goals using only visual observations, focusing on multi-step 3D box rearrangement tasks. Existing approaches typically rely on symbolic planners with brittle relational grounding of states and goals, or on direct action-sequence generation from 2D vision-language models (VLMs). Both approaches struggle with reasoning over many objects, rich 3D geometry, and implicit semantic constraints. Recent advances in 3D VLMs demonstrate strong grounding of natural-language referents to 3D segmentation masks, suggesting the potential for more general planning capabilities. We extend existing 3D grounding models and propose Reactive Action Mask Planner (RAMP-3D), which formulates long-horizon planning as sequential reactive prediction of paired 3D masks: a "which-object" mask indicating what to pick and a "which-target-region" mask specifying where to place it. The resulting system processes RGB-D observations and natural-language task specifications to reactively generate multi-step pick-and-place actions for 3D box rearrangement. We conduct experiments across 11 task variants in warehouse-style environments with 1-30 boxes and diverse natural-language constraints. RAMP-3D achieves 79.5% success rate on long-horizon rearrangement tasks and significantly outperforms 2D VLM-based baselines, establishing mask-based reactive policies as a promising alternative to symbolic pipelines for long-horizon planning.
title Grounding Vision and Language to 3D Masks for Long-Horizon Box Rearrangement
topic Artificial Intelligence
Robotics
url https://arxiv.org/abs/2603.23676