Saved in:
Bibliographic Details
Main Authors: Chen, Haonan, Li, Junxiao, Wu, Ruihai, Liu, Yiwei, Hou, Yiwen, Xu, Zhixuan, Guo, Jingxiang, Gao, Chongkai, Wei, Zhenyu, Xu, Shensi, Huang, Jiaqi, Shao, Lin
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.08372
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913987016261632
author Chen, Haonan
Li, Junxiao
Wu, Ruihai
Liu, Yiwei
Hou, Yiwen
Xu, Zhixuan
Guo, Jingxiang
Gao, Chongkai
Wei, Zhenyu
Xu, Shensi
Huang, Jiaqi
Shao, Lin
author_facet Chen, Haonan
Li, Junxiao
Wu, Ruihai
Liu, Yiwei
Hou, Yiwen
Xu, Zhixuan
Guo, Jingxiang
Gao, Chongkai
Wei, Zhenyu
Xu, Shensi
Huang, Jiaqi
Shao, Lin
contents Garment folding is a common yet challenging task in robotic manipulation. The deformability of garments leads to a vast state space and complex dynamics, which complicates precise and fine-grained manipulation. Previous approaches often rely on predefined key points or demonstrations, limiting their generalization across diverse garment categories. This paper presents a framework, MetaFold, that disentangles task planning from action prediction, learning each independently to enhance model generalization. It employs language-guided point cloud trajectory generation for task planning and a low-level foundation model for action prediction. This structure facilitates multi-category learning, enabling the model to adapt flexibly to various user instructions and folding tasks. Experimental results demonstrate the superiority of our proposed framework. Supplementary materials are available on our website: https://meta-fold.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08372
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MetaFold: Language-Guided Multi-Category Garment Folding Framework via Trajectory Generation and Foundation Model
Chen, Haonan
Li, Junxiao
Wu, Ruihai
Liu, Yiwei
Hou, Yiwen
Xu, Zhixuan
Guo, Jingxiang
Gao, Chongkai
Wei, Zhenyu
Xu, Shensi
Huang, Jiaqi
Shao, Lin
Robotics
Garment folding is a common yet challenging task in robotic manipulation. The deformability of garments leads to a vast state space and complex dynamics, which complicates precise and fine-grained manipulation. Previous approaches often rely on predefined key points or demonstrations, limiting their generalization across diverse garment categories. This paper presents a framework, MetaFold, that disentangles task planning from action prediction, learning each independently to enhance model generalization. It employs language-guided point cloud trajectory generation for task planning and a low-level foundation model for action prediction. This structure facilitates multi-category learning, enabling the model to adapt flexibly to various user instructions and folding tasks. Experimental results demonstrate the superiority of our proposed framework. Supplementary materials are available on our website: https://meta-fold.github.io/.
title MetaFold: Language-Guided Multi-Category Garment Folding Framework via Trajectory Generation and Foundation Model
topic Robotics
url https://arxiv.org/abs/2503.08372