Decomposed Object Manipulation via Dual-Actor Policy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Bin, Jiang, Jian-Jian, Li, Zhuohao, Wu, Xiao-Ming, He, Yi-Xiang, Yang, YiHan, Liu, Shengbang, Zheng, Wei-Shi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908714196271104
author Fan, Bin
Jiang, Jian-Jian
Li, Zhuohao
Wu, Xiao-Ming
He, Yi-Xiang
Yang, YiHan
Liu, Shengbang
Zheng, Wei-Shi
author_facet Fan, Bin
Jiang, Jian-Jian
Li, Zhuohao
Wu, Xiao-Ming
He, Yi-Xiang
Yang, YiHan
Liu, Shengbang
Zheng, Wei-Shi
contents Object manipulation, which focuses on learning to perform tasks on similar parts across different types of objects, can be divided into an approaching stage and a manipulation stage. However, previous works often ignore this characteristic of the task and rely on a single policy to directly learn the whole process of object manipulation. To address this problem, we propose a novel Dual-Actor Policy, termed DAP, which explicitly considers different stages and leverages heterogeneous visual priors to enhance each stage. Specifically, we introduce an affordance-based actor to locate the functional part in the manipulation task, thereby improving the approaching process. Following this, we propose a motion flow-based actor to capture the movement of the component, facilitating the manipulation process. Finally, we introduce a decision maker to determine the current stage of DAP and select the corresponding actor. Moreover, existing object manipulation datasets contain few objects and lack the visual priors needed to support training. To address this, we construct a simulated dataset, the Dual-Prior Object Manipulation Dataset, which combines the two visual priors and includes seven tasks, including two challenging long-term, multi-stage tasks. Experimental results on our dataset, the RoboTwin benchmark and real-world scenarios illustrate that our method consistently outperforms the SOTA method by 5.55%, 14.7% and 10.4% on average respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05129
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Decomposed Object Manipulation via Dual-Actor Policy
Fan, Bin
Jiang, Jian-Jian
Li, Zhuohao
Wu, Xiao-Ming
He, Yi-Xiang
Yang, YiHan
Liu, Shengbang
Zheng, Wei-Shi
Robotics
Object manipulation, which focuses on learning to perform tasks on similar parts across different types of objects, can be divided into an approaching stage and a manipulation stage. However, previous works often ignore this characteristic of the task and rely on a single policy to directly learn the whole process of object manipulation. To address this problem, we propose a novel Dual-Actor Policy, termed DAP, which explicitly considers different stages and leverages heterogeneous visual priors to enhance each stage. Specifically, we introduce an affordance-based actor to locate the functional part in the manipulation task, thereby improving the approaching process. Following this, we propose a motion flow-based actor to capture the movement of the component, facilitating the manipulation process. Finally, we introduce a decision maker to determine the current stage of DAP and select the corresponding actor. Moreover, existing object manipulation datasets contain few objects and lack the visual priors needed to support training. To address this, we construct a simulated dataset, the Dual-Prior Object Manipulation Dataset, which combines the two visual priors and includes seven tasks, including two challenging long-term, multi-stage tasks. Experimental results on our dataset, the RoboTwin benchmark and real-world scenarios illustrate that our method consistently outperforms the SOTA method by 5.55%, 14.7% and 10.4% on average respectively.
title Decomposed Object Manipulation via Dual-Actor Policy
topic Robotics
url https://arxiv.org/abs/2511.05129