AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Mingyang, Xu, Haofan, Sun, Haowen, Chen, Xinzhe, Ren, Sihua, Huang, Liqi, Sui, Xinyang, Miao, Chenyang, Ye, Jiawei, Cui, Qiongjie, Liu, Zeyang, Chen, Xingyu, Lan, Xuguang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915999725387776
author Li, Mingyang
Xu, Haofan
Sun, Haowen
Chen, Xinzhe
Ren, Sihua
Huang, Liqi
Sui, Xinyang
Miao, Chenyang
Ye, Jiawei
Cui, Qiongjie
Liu, Zeyang
Chen, Xingyu
Lan, Xuguang
author_facet Li, Mingyang
Xu, Haofan
Sun, Haowen
Chen, Xinzhe
Ren, Sihua
Huang, Liqi
Sui, Xinyang
Miao, Chenyang
Ye, Jiawei
Cui, Qiongjie
Liu, Zeyang
Chen, Xingyu
Lan, Xuguang
contents Many everyday robot manipulation skills are affordance-dependent, with success determined by whether the robot contacts the functional object region required by the subsequent action. Current simulation data generators obtain contacts from generic grasp estimators or per-object manual contact annotations, but generic estimators rank stable grasps without task semantics and often select contacts that are misaligned with the downstream action, while manual contact annotations must be rewritten for each new object and task. To solve these challenges, we introduce AffordSim, a scalable data generator and benchmark that integrates open-vocabulary 3D affordance prediction into simulation-based trajectory generation. Given a natural-language task description, AffordSim synthesizes a task-relevant scene, emits affordance queries, grounds them on object surfaces, samples region-conditioned grasps, and selects executable candidates with motion planning. It further randomizes object pose, texture, lighting, image noise, and cross-viewpoint backgrounds for sim-to-real transfer. We instantiate AffordSim as a 50-task benchmark across diverse manipulation skills, five robot embodiments, and 500+ rigid and articulated objects. AffordSim achieves 93% of the trajectory collection success rate of manual contact annotations on affordance-critical tasks and 89% on hard composite tasks. Vision-language-action policies trained on AffordSim data transfer zero-shot to a real Franka FR3, reaching 24% average success.
format Preprint
id arxiv_https___arxiv_org_abs_2604_11674
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation
Li, Mingyang
Xu, Haofan
Sun, Haowen
Chen, Xinzhe
Ren, Sihua
Huang, Liqi
Sui, Xinyang
Miao, Chenyang
Ye, Jiawei
Cui, Qiongjie
Liu, Zeyang
Chen, Xingyu
Lan, Xuguang
Robotics
Artificial Intelligence
Many everyday robot manipulation skills are affordance-dependent, with success determined by whether the robot contacts the functional object region required by the subsequent action. Current simulation data generators obtain contacts from generic grasp estimators or per-object manual contact annotations, but generic estimators rank stable grasps without task semantics and often select contacts that are misaligned with the downstream action, while manual contact annotations must be rewritten for each new object and task. To solve these challenges, we introduce AffordSim, a scalable data generator and benchmark that integrates open-vocabulary 3D affordance prediction into simulation-based trajectory generation. Given a natural-language task description, AffordSim synthesizes a task-relevant scene, emits affordance queries, grounds them on object surfaces, samples region-conditioned grasps, and selects executable candidates with motion planning. It further randomizes object pose, texture, lighting, image noise, and cross-viewpoint backgrounds for sim-to-real transfer. We instantiate AffordSim as a 50-task benchmark across diverse manipulation skills, five robot embodiments, and 500+ rigid and articulated objects. AffordSim achieves 93% of the trajectory collection success rate of manual contact annotations on affordance-critical tasks and 89% on hard composite tasks. Vision-language-action policies trained on AffordSim data transfer zero-shot to a real Franka FR3, reaching 24% average success.
title AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2604.11674