Vision in Action: Learning Active Perception from Human Demonstrations
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918063150989312 |
|---|---|
| author | Xiong, Haoyu Xu, Xiaomeng Wu, Jimmy Hou, Yifan Bohg, Jeannette Song, Shuran |
| author_facet | Xiong, Haoyu Xu, Xiaomeng Wu, Jimmy Hou, Yifan Bohg, Jeannette Song, Shuran |
| contents | We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the hardware side, ViA employs a simple yet effective 6-DoF robotic neck to enable flexible, human-like head movements. To capture human active perception strategies, we design a VR-based teleoperation interface that creates a shared observation space between the robot and the human operator. To mitigate VR motion sickness caused by latency in the robot's physical movements, the interface uses an intermediate 3D scene representation, enabling real-time view rendering on the operator side while asynchronously updating the scene with the robot's latest observations. Together, these design elements enable the learning of robust visuomotor policies for three complex, multi-stage bimanual manipulation tasks involving visual occlusions, significantly outperforming baseline systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_15666 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Vision in Action: Learning Active Perception from Human Demonstrations Xiong, Haoyu Xu, Xiaomeng Wu, Jimmy Hou, Yifan Bohg, Jeannette Song, Shuran Robotics We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the hardware side, ViA employs a simple yet effective 6-DoF robotic neck to enable flexible, human-like head movements. To capture human active perception strategies, we design a VR-based teleoperation interface that creates a shared observation space between the robot and the human operator. To mitigate VR motion sickness caused by latency in the robot's physical movements, the interface uses an intermediate 3D scene representation, enabling real-time view rendering on the operator side while asynchronously updating the scene with the robot's latest observations. Together, these design elements enable the learning of robust visuomotor policies for three complex, multi-stage bimanual manipulation tasks involving visual occlusions, significantly outperforming baseline systems. |
| title | Vision in Action: Learning Active Perception from Human Demonstrations |
| topic | Robotics |
| url | https://arxiv.org/abs/2506.15666 |