UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Yihe, Huang, Wenlong, Wang, Yingke, Li, Chengshu, Yuan, Roy, Zhang, Ruohan, Wu, Jiajun, Fei-Fei, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914004889239552
author Tang, Yihe
Huang, Wenlong
Wang, Yingke
Li, Chengshu
Yuan, Roy
Zhang, Ruohan
Wu, Jiajun
Fei-Fei, Li
author_facet Tang, Yihe
Huang, Wenlong
Wang, Yingke
Li, Chengshu
Yuan, Roy
Zhang, Ruohan
Wu, Jiajun
Fei-Fei, Li
contents Understanding fine-grained object affordances is imperative for robots to manipulate objects in unstructured environments given open-ended task instructions. However, existing methods of visual affordance predictions often rely on manually annotated data or conditions only on a predefined set of tasks. We introduce UAD (Unsupervised Affordance Distillation), a method for distilling affordance knowledge from foundation models into a task-conditioned affordance model without any manual annotations. By leveraging the complementary strengths of large vision models and vision-language models, UAD automatically annotates a large-scale dataset with detailed $<$instruction, visual affordance$>$ pairs. Training only a lightweight task-conditioned decoder atop frozen features, UAD exhibits notable generalization to in-the-wild robotic scenes and to various human activities, despite only being trained on rendered objects in simulation. Using affordance provided by UAD as the observation space, we show an imitation learning policy that demonstrates promising generalization to unseen object instances, object categories, and even variations in task instructions after training on as few as 10 demonstrations. Project website: https://unsup-affordance.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2506_09284
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
Tang, Yihe
Huang, Wenlong
Wang, Yingke
Li, Chengshu
Yuan, Roy
Zhang, Ruohan
Wu, Jiajun
Fei-Fei, Li
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Understanding fine-grained object affordances is imperative for robots to manipulate objects in unstructured environments given open-ended task instructions. However, existing methods of visual affordance predictions often rely on manually annotated data or conditions only on a predefined set of tasks. We introduce UAD (Unsupervised Affordance Distillation), a method for distilling affordance knowledge from foundation models into a task-conditioned affordance model without any manual annotations. By leveraging the complementary strengths of large vision models and vision-language models, UAD automatically annotates a large-scale dataset with detailed $<$instruction, visual affordance$>$ pairs. Training only a lightweight task-conditioned decoder atop frozen features, UAD exhibits notable generalization to in-the-wild robotic scenes and to various human activities, despite only being trained on rendered objects in simulation. Using affordance provided by UAD as the observation space, we show an imitation learning policy that demonstrates promising generalization to unseen object instances, object categories, and even variations in task instructions after training on as few as 10 demonstrations. Project website: https://unsup-affordance.github.io/
title UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.09284