Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Changshi, Xu, Haichuan, Gu, Ningquan, Wang, Zhipeng, Cheng, Bin, Zhang, Pengpeng, Dong, Yanchao, Hayashibe, Mitsuhiro, Zhou, Yanmin, He, Bin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916930130018304
author Zhou, Changshi
Xu, Haichuan
Gu, Ningquan
Wang, Zhipeng
Cheng, Bin
Zhang, Pengpeng
Dong, Yanchao
Hayashibe, Mitsuhiro
Zhou, Yanmin
He, Bin
author_facet Zhou, Changshi
Xu, Haichuan
Gu, Ningquan
Wang, Zhipeng
Cheng, Bin
Zhang, Pengpeng
Dong, Yanchao
Hayashibe, Mitsuhiro
Zhou, Yanmin
He, Bin
contents Language-guided long-horizon manipulation of deformable objects presents significant challenges due to high degrees of freedom, complex dynamics, and the need for accurate vision-language grounding. In this work, we focus on multi-step cloth folding, a representative deformable-object manipulation task that requires both structured long-horizon planning and fine-grained visual perception. To this end, we propose a unified framework that integrates a Large Language Model (LLM)-based planner, a Vision-Language Model (VLM)-based perception system, and a task execution module. Specifically, the LLM-based planner decomposes high-level language instructions into low-level action primitives, bridging the semantic-execution gap, aligning perception with action, and enhancing generalization. The VLM-based perception module employs a SigLIP2-driven architecture with a bidirectional cross-attention fusion mechanism and weight-decomposed low-rank adaptation (DoRA) fine-tuning to achieve language-conditioned fine-grained visual grounding. Experiments in both simulation and real-world settings demonstrate the method's effectiveness. In simulation, it outperforms state-of-the-art baselines by 2.23, 1.87, and 33.3 on seen instructions, unseen instructions, and unseen tasks, respectively. On a real robot, it robustly executes multi-step folding sequences from language instructions across diverse cloth materials and configurations, demonstrating strong generalization in practical scenarios. Project page: https://language-guided.netlify.app/
format Preprint
id arxiv_https___arxiv_org_abs_2509_02324
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception
Zhou, Changshi
Xu, Haichuan
Gu, Ningquan
Wang, Zhipeng
Cheng, Bin
Zhang, Pengpeng
Dong, Yanchao
Hayashibe, Mitsuhiro
Zhou, Yanmin
He, Bin
Robotics
Language-guided long-horizon manipulation of deformable objects presents significant challenges due to high degrees of freedom, complex dynamics, and the need for accurate vision-language grounding. In this work, we focus on multi-step cloth folding, a representative deformable-object manipulation task that requires both structured long-horizon planning and fine-grained visual perception. To this end, we propose a unified framework that integrates a Large Language Model (LLM)-based planner, a Vision-Language Model (VLM)-based perception system, and a task execution module. Specifically, the LLM-based planner decomposes high-level language instructions into low-level action primitives, bridging the semantic-execution gap, aligning perception with action, and enhancing generalization. The VLM-based perception module employs a SigLIP2-driven architecture with a bidirectional cross-attention fusion mechanism and weight-decomposed low-rank adaptation (DoRA) fine-tuning to achieve language-conditioned fine-grained visual grounding. Experiments in both simulation and real-world settings demonstrate the method's effectiveness. In simulation, it outperforms state-of-the-art baselines by 2.23, 1.87, and 33.3 on seen instructions, unseen instructions, and unseen tasks, respectively. On a real robot, it robustly executes multi-step folding sequences from language instructions across diverse cloth materials and configurations, demonstrating strong generalization in practical scenarios. Project page: https://language-guided.netlify.app/
title Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception
topic Robotics
url https://arxiv.org/abs/2509.02324