A0: An Affordance-Aware Hierarchical Model for General Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917210573766656 |
|---|---|
| author | Xu, Rongtao Zhang, Jian Guo, Minghao Wen, Youpeng Yang, Haoting Lin, Min Huang, Jianzheng Li, Zhe Zhang, Kaidong Wang, Liqiong Kuang, Yuxuan Cao, Meng Zheng, Feng Liang, Xiaodan |
| author_facet | Xu, Rongtao Zhang, Jian Guo, Minghao Wen, Youpeng Yang, Haoting Lin, Min Huang, Jianzheng Li, Zhe Zhang, Kaidong Wang, Liqiong Kuang, Yuxuan Cao, Meng Zheng, Feng Liang, Xiaodan |
| contents | Robotic manipulation faces critical challenges in understanding spatial affordances--the "where" and "how" of object interactions--essential for complex manipulation tasks like wiping a board or stacking objects. Existing methods, including modular-based and end-to-end approaches, often lack robust spatial reasoning capabilities. Unlike recent point-based and flow-based affordance methods that focus on dense spatial representations or trajectory modeling, we propose A0, a hierarchical affordance-aware diffusion model that decomposes manipulation tasks into high-level spatial affordance understanding and low-level action execution. A0 leverages the Embodiment-Agnostic Affordance Representation, which captures object-centric spatial affordances by predicting contact points and post-contact trajectories. A0 is pre-trained on 1 million contact points data and fine-tuned on annotated trajectories, enabling generalization across platforms. Key components include Position Offset Attention for motion-aware feature extraction and a Spatial Information Aggregation Layer for precise coordinate mapping. The model's output is executed by the action execution module. Experiments on multiple robotic systems (Franka, Kinova, Realman, and Dobot) demonstrate A0's superior performance in complex tasks, showcasing its efficiency, flexibility, and real-world applicability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_12636 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A0: An Affordance-Aware Hierarchical Model for General Robotic Manipulation Xu, Rongtao Zhang, Jian Guo, Minghao Wen, Youpeng Yang, Haoting Lin, Min Huang, Jianzheng Li, Zhe Zhang, Kaidong Wang, Liqiong Kuang, Yuxuan Cao, Meng Zheng, Feng Liang, Xiaodan Robotics Robotic manipulation faces critical challenges in understanding spatial affordances--the "where" and "how" of object interactions--essential for complex manipulation tasks like wiping a board or stacking objects. Existing methods, including modular-based and end-to-end approaches, often lack robust spatial reasoning capabilities. Unlike recent point-based and flow-based affordance methods that focus on dense spatial representations or trajectory modeling, we propose A0, a hierarchical affordance-aware diffusion model that decomposes manipulation tasks into high-level spatial affordance understanding and low-level action execution. A0 leverages the Embodiment-Agnostic Affordance Representation, which captures object-centric spatial affordances by predicting contact points and post-contact trajectories. A0 is pre-trained on 1 million contact points data and fine-tuned on annotated trajectories, enabling generalization across platforms. Key components include Position Offset Attention for motion-aware feature extraction and a Spatial Information Aggregation Layer for precise coordinate mapping. The model's output is executed by the action execution module. Experiments on multiple robotic systems (Franka, Kinova, Realman, and Dobot) demonstrate A0's superior performance in complex tasks, showcasing its efficiency, flexibility, and real-world applicability. |
| title | A0: An Affordance-Aware Hierarchical Model for General Robotic Manipulation |
| topic | Robotics |
| url | https://arxiv.org/abs/2504.12636 |