AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Jiawei, Hu, Kaizhe, Huang, Yingqian, Ju, Yuanchen, Xue, Zhengrong, Xu, Huazhe
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911735567351808
author Zhang, Jiawei
Hu, Kaizhe
Huang, Yingqian
Ju, Yuanchen
Xue, Zhengrong
Xu, Huazhe
author_facet Zhang, Jiawei
Hu, Kaizhe
Huang, Yingqian
Ju, Yuanchen
Xue, Zhengrong
Xu, Huazhe
contents Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations due to limited data diversity. Leveraging powerful 3D generative models and vision foundation models (VFMs), the proposed AffordGen framework overcomes this limitation by utilizing the semantic correspondence of meaningful keypoints across large-scale 3D meshes to generate new robot manipulation trajectories. This large-scale, affordance-aware dataset is then used to train a robust, closed-loop visuomotor policy, combining the semantic generalizability of affordances with the reactive robustness of end-to-end learning. Experiments in simulation and the real world show that policies trained with AffordGen achieve high success rates and enable zero-shot generalization to truly unseen objects, significantly improving data efficiency in robot learning. Project Page: https://jiaweiz9.github.io/AffordGen-release/
format Preprint
id arxiv_https___arxiv_org_abs_2604_10579
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence
Zhang, Jiawei
Hu, Kaizhe
Huang, Yingqian
Ju, Yuanchen
Xue, Zhengrong
Xu, Huazhe
Robotics
Artificial Intelligence
Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations due to limited data diversity. Leveraging powerful 3D generative models and vision foundation models (VFMs), the proposed AffordGen framework overcomes this limitation by utilizing the semantic correspondence of meaningful keypoints across large-scale 3D meshes to generate new robot manipulation trajectories. This large-scale, affordance-aware dataset is then used to train a robust, closed-loop visuomotor policy, combining the semantic generalizability of affordances with the reactive robustness of end-to-end learning. Experiments in simulation and the real world show that policies trained with AffordGen achieve high success rates and enable zero-shot generalization to truly unseen objects, significantly improving data efficiency in robot learning. Project Page: https://jiaweiz9.github.io/AffordGen-release/
title AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2604.10579