Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918349021118464 |
|---|---|
| author | Hu, Jianshu Wang, Lidi Li, Shujia Jiang, Yunpeng Li, Xiao Weng, Paul Ban, Yutong |
| author_facet | Hu, Jianshu Wang, Lidi Li, Shujia Jiang, Yunpeng Li, Xiao Weng, Paul Ban, Yutong |
| contents | Hierarchical coarse-to-fine policy, where a coarse branch predicts a region of interest to guide a fine-grained action predictor, has demonstrated significant potential in robotic 3D manipulation tasks by especially enhancing sample efficiency and enabling more precise manipulation. However, even augmented with pre-trained models, these hierarchical policies still suffer from generalization issues. To enhance generalization to novel instructions and environment variations, we propose Coarse-to-fine Language-Aligned manipulation Policy (CLAP), a framework that integrates three key components: 1) task decomposition, 2) VLM fine-tuning for 3D keypoint prediction, and 3) 3D-aware representation. Through comprehensive experiments in simulation and on a real robot, we demonstrate its superior generalization capability. Specifically, on GemBench, a benchmark designed for evaluating generalization, our approach achieves a 12\% higher average success rate than the SOTA method while using only 1/5 of the training trajectories. In real-world experiments, our policy, trained on only 10 demonstrations, successfully generalizes to novel instructions and environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_23575 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints Hu, Jianshu Wang, Lidi Li, Shujia Jiang, Yunpeng Li, Xiao Weng, Paul Ban, Yutong Robotics Hierarchical coarse-to-fine policy, where a coarse branch predicts a region of interest to guide a fine-grained action predictor, has demonstrated significant potential in robotic 3D manipulation tasks by especially enhancing sample efficiency and enabling more precise manipulation. However, even augmented with pre-trained models, these hierarchical policies still suffer from generalization issues. To enhance generalization to novel instructions and environment variations, we propose Coarse-to-fine Language-Aligned manipulation Policy (CLAP), a framework that integrates three key components: 1) task decomposition, 2) VLM fine-tuning for 3D keypoint prediction, and 3) 3D-aware representation. Through comprehensive experiments in simulation and on a real robot, we demonstrate its superior generalization capability. Specifically, on GemBench, a benchmark designed for evaluating generalization, our approach achieves a 12\% higher average success rate than the SOTA method while using only 1/5 of the training trajectories. In real-world experiments, our policy, trained on only 10 demonstrations, successfully generalizes to novel instructions and environments. |
| title | Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints |
| topic | Robotics |
| url | https://arxiv.org/abs/2509.23575 |