Vision in Action: Learning Active Perception from Human Demonstrations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xiong, Haoyu, Xu, Xiaomeng, Wu, Jimmy, Hou, Yifan, Bohg, Jeannette, Song, Shuran
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918063150989312
author Xiong, Haoyu
Xu, Xiaomeng
Wu, Jimmy
Hou, Yifan
Bohg, Jeannette
Song, Shuran
author_facet Xiong, Haoyu
Xu, Xiaomeng
Wu, Jimmy
Hou, Yifan
Bohg, Jeannette
Song, Shuran
contents We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the hardware side, ViA employs a simple yet effective 6-DoF robotic neck to enable flexible, human-like head movements. To capture human active perception strategies, we design a VR-based teleoperation interface that creates a shared observation space between the robot and the human operator. To mitigate VR motion sickness caused by latency in the robot's physical movements, the interface uses an intermediate 3D scene representation, enabling real-time view rendering on the operator side while asynchronously updating the scene with the robot's latest observations. Together, these design elements enable the learning of robust visuomotor policies for three complex, multi-stage bimanual manipulation tasks involving visual occlusions, significantly outperforming baseline systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15666
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision in Action: Learning Active Perception from Human Demonstrations
Xiong, Haoyu
Xu, Xiaomeng
Wu, Jimmy
Hou, Yifan
Bohg, Jeannette
Song, Shuran
Robotics
We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the hardware side, ViA employs a simple yet effective 6-DoF robotic neck to enable flexible, human-like head movements. To capture human active perception strategies, we design a VR-based teleoperation interface that creates a shared observation space between the robot and the human operator. To mitigate VR motion sickness caused by latency in the robot's physical movements, the interface uses an intermediate 3D scene representation, enabling real-time view rendering on the operator side while asynchronously updating the scene with the robot's latest observations. Together, these design elements enable the learning of robust visuomotor policies for three complex, multi-stage bimanual manipulation tasks involving visual occlusions, significantly outperforming baseline systems.
title Vision in Action: Learning Active Perception from Human Demonstrations
topic Robotics
url https://arxiv.org/abs/2506.15666