Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Guokang, Li, Hang, Zhang, Shuyuan, Guo, Di, Liu, Yanhong, Liu, Huaping
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909488830742528
author Wang, Guokang
Li, Hang
Zhang, Shuyuan
Guo, Di
Liu, Yanhong
Liu, Huaping
author_facet Wang, Guokang
Li, Hang
Zhang, Shuyuan
Guo, Di
Liu, Yanhong
Liu, Huaping
contents In real-world scenarios, many robotic manipulation tasks are hindered by occlusions and limited fields of view, posing significant challenges for passive observation-based models that rely on fixed or wrist-mounted cameras. In this paper, we investigate the problem of robotic manipulation under limited visual observation and propose a task-driven asynchronous active vision-action model.Our model serially connects a camera Next-Best-View (NBV) policy with a gripper Next-Best Pose (NBP) policy, and trains them in a sensor-motor coordination framework using few-shot reinforcement learning. This approach allows the agent to adjust a third-person camera to actively observe the environment based on the task goal, and subsequently infer the appropriate manipulation actions.We trained and evaluated our model on 8 viewpoint-constrained tasks in RLBench. The results demonstrate that our model consistently outperforms baseline algorithms, showcasing its effectiveness in handling visual constraints in manipulation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2409_14891
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
Wang, Guokang
Li, Hang
Zhang, Shuyuan
Guo, Di
Liu, Yanhong
Liu, Huaping
Robotics
Computer Vision and Pattern Recognition
In real-world scenarios, many robotic manipulation tasks are hindered by occlusions and limited fields of view, posing significant challenges for passive observation-based models that rely on fixed or wrist-mounted cameras. In this paper, we investigate the problem of robotic manipulation under limited visual observation and propose a task-driven asynchronous active vision-action model.Our model serially connects a camera Next-Best-View (NBV) policy with a gripper Next-Best Pose (NBP) policy, and trains them in a sensor-motor coordination framework using few-shot reinforcement learning. This approach allows the agent to adjust a third-person camera to actively observe the environment based on the task goal, and subsequently infer the appropriate manipulation actions.We trained and evaluated our model on 8 viewpoint-constrained tasks in RLBench. The results demonstrate that our model consistently outperforms baseline algorithms, showcasing its effectiveness in handling visual constraints in manipulation tasks.
title Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.14891