Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909488830742528 |
|---|---|
| author | Wang, Guokang Li, Hang Zhang, Shuyuan Guo, Di Liu, Yanhong Liu, Huaping |
| author_facet | Wang, Guokang Li, Hang Zhang, Shuyuan Guo, Di Liu, Yanhong Liu, Huaping |
| contents | In real-world scenarios, many robotic manipulation tasks are hindered by occlusions and limited fields of view, posing significant challenges for passive observation-based models that rely on fixed or wrist-mounted cameras. In this paper, we investigate the problem of robotic manipulation under limited visual observation and propose a task-driven asynchronous active vision-action model.Our model serially connects a camera Next-Best-View (NBV) policy with a gripper Next-Best Pose (NBP) policy, and trains them in a sensor-motor coordination framework using few-shot reinforcement learning. This approach allows the agent to adjust a third-person camera to actively observe the environment based on the task goal, and subsequently infer the appropriate manipulation actions.We trained and evaluated our model on 8 viewpoint-constrained tasks in RLBench. The results demonstrate that our model consistently outperforms baseline algorithms, showcasing its effectiveness in handling visual constraints in manipulation tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_14891 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation Wang, Guokang Li, Hang Zhang, Shuyuan Guo, Di Liu, Yanhong Liu, Huaping Robotics Computer Vision and Pattern Recognition In real-world scenarios, many robotic manipulation tasks are hindered by occlusions and limited fields of view, posing significant challenges for passive observation-based models that rely on fixed or wrist-mounted cameras. In this paper, we investigate the problem of robotic manipulation under limited visual observation and propose a task-driven asynchronous active vision-action model.Our model serially connects a camera Next-Best-View (NBV) policy with a gripper Next-Best Pose (NBP) policy, and trains them in a sensor-motor coordination framework using few-shot reinforcement learning. This approach allows the agent to adjust a third-person camera to actively observe the environment based on the task goal, and subsequently infer the appropriate manipulation actions.We trained and evaluated our model on 8 viewpoint-constrained tasks in RLBench. The results demonstrate that our model consistently outperforms baseline algorithms, showcasing its effectiveness in handling visual constraints in manipulation tasks. |
| title | Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2409.14891 |