ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Zichao, Tang, Chen, Munje, Michael J., Zhu, Yifeng, Liu, Alex, Liu, Shuijing, Warnell, Garrett, Stone, Peter, Biswas, Joydeep
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914051109421056
author Hu, Zichao
Tang, Chen
Munje, Michael J.
Zhu, Yifeng
Liu, Alex
Liu, Shuijing
Warnell, Garrett
Stone, Peter
Biswas, Joydeep
author_facet Hu, Zichao
Tang, Chen
Munje, Michael J.
Zhu, Yifeng
Liu, Alex
Liu, Shuijing
Warnell, Garrett
Stone, Peter
Biswas, Joydeep
contents This paper considers the problem of enabling robots to navigate dynamic environments while following instructions. The challenge lies in the combinatorial nature of instruction specifications: each instruction can include multiple specifications, and the number of possible specification combinations grows exponentially as the robot's skill set expands. For example, "overtake the pedestrian while staying on the right side of the road" consists of two specifications: "overtake the pedestrian" and "walk on the right side of the road." To tackle this challenge, we propose ComposableNav, based on the intuition that following an instruction involves independently satisfying its constituent specifications, each corresponding to a distinct motion primitive. Using diffusion models, ComposableNav learns each primitive separately, then composes them in parallel at deployment time to satisfy novel combinations of specifications unseen in training. Additionally, to avoid the onerous need for demonstrations of individual motion primitives, we propose a two-stage training procedure: (1) supervised pre-training to learn a base diffusion model for dynamic navigation, and (2) reinforcement learning fine-tuning that molds the base model into different motion primitives. Through simulation and real-world experiments, we show that ComposableNav enables robots to follow instructions by generating trajectories that satisfy diverse and unseen combinations of specifications, significantly outperforming both non-compositional VLM-based policies and costmap composing baselines. Videos and additional materials can be found on the project page: https://amrl.cs.utexas.edu/ComposableNav/
format Preprint
id arxiv_https___arxiv_org_abs_2509_17941
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion
Hu, Zichao
Tang, Chen
Munje, Michael J.
Zhu, Yifeng
Liu, Alex
Liu, Shuijing
Warnell, Garrett
Stone, Peter
Biswas, Joydeep
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
This paper considers the problem of enabling robots to navigate dynamic environments while following instructions. The challenge lies in the combinatorial nature of instruction specifications: each instruction can include multiple specifications, and the number of possible specification combinations grows exponentially as the robot's skill set expands. For example, "overtake the pedestrian while staying on the right side of the road" consists of two specifications: "overtake the pedestrian" and "walk on the right side of the road." To tackle this challenge, we propose ComposableNav, based on the intuition that following an instruction involves independently satisfying its constituent specifications, each corresponding to a distinct motion primitive. Using diffusion models, ComposableNav learns each primitive separately, then composes them in parallel at deployment time to satisfy novel combinations of specifications unseen in training. Additionally, to avoid the onerous need for demonstrations of individual motion primitives, we propose a two-stage training procedure: (1) supervised pre-training to learn a base diffusion model for dynamic navigation, and (2) reinforcement learning fine-tuning that molds the base model into different motion primitives. Through simulation and real-world experiments, we show that ComposableNav enables robots to follow instructions by generating trajectories that satisfy diverse and unseen combinations of specifications, significantly outperforming both non-compositional VLM-based policies and costmap composing baselines. Videos and additional materials can be found on the project page: https://amrl.cs.utexas.edu/ComposableNav/
title ComposableNav: Instruction-Following Navigation in Dynamic Environments via Composable Diffusion
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.17941