Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Zhecheng, Wei, Tianming, Cheng, Shuiqi, Zhang, Gu, Chen, Yuanpei, Xu, Huazhe
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916449291862016
author Yuan, Zhecheng
Wei, Tianming
Cheng, Shuiqi
Zhang, Gu
Chen, Yuanpei
Xu, Huazhe
author_facet Yuan, Zhecheng
Wei, Tianming
Cheng, Shuiqi
Zhang, Gu
Chen, Yuanpei
Xu, Huazhe
contents Can we endow visuomotor robots with generalization capabilities to operate in diverse open-world scenarios? In this paper, we propose \textbf{Maniwhere}, a generalizable framework tailored for visual reinforcement learning, enabling the trained robot policies to generalize across a combination of multiple visual disturbance types. Specifically, we introduce a multi-view representation learning approach fused with Spatial Transformer Network (STN) module to capture shared semantic information and correspondences among different viewpoints. In addition, we employ a curriculum-based randomization and augmentation approach to stabilize the RL training process and strengthen the visual generalization ability. To exhibit the effectiveness of Maniwhere, we meticulously design 8 tasks encompassing articulate objects, bi-manual, and dexterous hand manipulation tasks, demonstrating Maniwhere's strong visual generalization and sim2real transfer abilities across 3 hardware platforms. Our experiments show that Maniwhere significantly outperforms existing state-of-the-art methods. Videos are provided at https://gemcollector.github.io/maniwhere/.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15815
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning
Yuan, Zhecheng
Wei, Tianming
Cheng, Shuiqi
Zhang, Gu
Chen, Yuanpei
Xu, Huazhe
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Can we endow visuomotor robots with generalization capabilities to operate in diverse open-world scenarios? In this paper, we propose \textbf{Maniwhere}, a generalizable framework tailored for visual reinforcement learning, enabling the trained robot policies to generalize across a combination of multiple visual disturbance types. Specifically, we introduce a multi-view representation learning approach fused with Spatial Transformer Network (STN) module to capture shared semantic information and correspondences among different viewpoints. In addition, we employ a curriculum-based randomization and augmentation approach to stabilize the RL training process and strengthen the visual generalization ability. To exhibit the effectiveness of Maniwhere, we meticulously design 8 tasks encompassing articulate objects, bi-manual, and dexterous hand manipulation tasks, demonstrating Maniwhere's strong visual generalization and sim2real transfer abilities across 3 hardware platforms. Our experiments show that Maniwhere significantly outperforms existing state-of-the-art methods. Videos are provided at https://gemcollector.github.io/maniwhere/.
title Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.15815