Saved in:
Bibliographic Details
Main Authors: Ye, Haoming, Xiao, Yunxiao, Lu, Cewu, Cai, Panpan
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.08537
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910016506691584
author Ye, Haoming
Xiao, Yunxiao
Lu, Cewu
Cai, Panpan
author_facet Ye, Haoming
Xiao, Yunxiao
Lu, Cewu
Cai, Panpan
contents Integration of VLM reasoning with symbolic planning has proven to be a promising approach to real-world robot task planning. Existing work like UniDomain effectively learns symbolic manipulation domains from real-world demonstrations, described in Planning Domain Definition Language (PDDL), and has successfully applied them to real-world tasks. These domains, however, are restricted to tabletop manipulation. We propose UniPlan, a vision-language task planning system for long-horizon mobile-manipulation in large-scale indoor environments, that unifies scene topology, visuals, and robot capabilities into a holistic PDDL representation. UniPlan programmatically extends learned tabletop domains from UniDomain to support navigation, door traversal, and bimanual coordination. It operates on a visual-topological map, comprising navigation landmarks anchored with scene images. Given a language instruction, UniPlan retrieves task-relevant nodes from the map and uses a VLM to ground the anchored image into task-relevant objects and their PDDL states; next, it reconnects these nodes to a compressed, densely-connected topological map, also represented in PDDL, with connectivity and costs derived from the original map; Finally, a mobile-manipulation plan is generated using off-the-shelf PDDL solvers. Evaluated on human-raised tasks in a large-scale map with real-world imagery, UniPlan significantly outperforms VLM and LLM+PDDL planning in success rate, plan quality, and computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08537
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UniPlan: Vision-Language Task Planning for Mobile Manipulation with Unified PDDL Formulation
Ye, Haoming
Xiao, Yunxiao
Lu, Cewu
Cai, Panpan
Robotics
Integration of VLM reasoning with symbolic planning has proven to be a promising approach to real-world robot task planning. Existing work like UniDomain effectively learns symbolic manipulation domains from real-world demonstrations, described in Planning Domain Definition Language (PDDL), and has successfully applied them to real-world tasks. These domains, however, are restricted to tabletop manipulation. We propose UniPlan, a vision-language task planning system for long-horizon mobile-manipulation in large-scale indoor environments, that unifies scene topology, visuals, and robot capabilities into a holistic PDDL representation. UniPlan programmatically extends learned tabletop domains from UniDomain to support navigation, door traversal, and bimanual coordination. It operates on a visual-topological map, comprising navigation landmarks anchored with scene images. Given a language instruction, UniPlan retrieves task-relevant nodes from the map and uses a VLM to ground the anchored image into task-relevant objects and their PDDL states; next, it reconnects these nodes to a compressed, densely-connected topological map, also represented in PDDL, with connectivity and costs derived from the original map; Finally, a mobile-manipulation plan is generated using off-the-shelf PDDL solvers. Evaluated on human-raised tasks in a large-scale map with real-world imagery, UniPlan significantly outperforms VLM and LLM+PDDL planning in success rate, plan quality, and computational efficiency.
title UniPlan: Vision-Language Task Planning for Mobile Manipulation with Unified PDDL Formulation
topic Robotics
url https://arxiv.org/abs/2602.08537