MV-UMI: A Scalable Multi-View Interface for Cross-Embodiment Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Rayyan, Omar, Abanes, John, Hafez, Mahmoud, Tzes, Anthony, Abu-Dakka, Fares
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914052211474432
author Rayyan, Omar
Abanes, John
Hafez, Mahmoud
Tzes, Anthony
Abu-Dakka, Fares
author_facet Rayyan, Omar
Abanes, John
Hafez, Mahmoud
Tzes, Anthony
Abu-Dakka, Fares
contents Recent advances in imitation learning have shown great promise for developing robust robot manipulation policies from demonstrations. However, this promise is contingent on the availability of diverse, high-quality datasets, which are not only challenging and costly to collect but are often constrained to a specific robot embodiment. Portable handheld grippers have recently emerged as intuitive and scalable alternatives to traditional robotic teleoperation methods for data collection. However, their reliance solely on first-person view wrist-mounted cameras often creates limitations in capturing sufficient scene contexts. In this paper, we present MV-UMI (Multi-View Universal Manipulation Interface), a framework that integrates a third-person perspective with the egocentric camera to overcome this limitation. This integration mitigates domain shifts between human demonstration and robot deployment, preserving the cross-embodiment advantages of handheld data-collection devices. Our experimental results, including an ablation study, demonstrate that our MV-UMI framework improves performance in sub-tasks requiring broad scene understanding by approximately 47% across 3 tasks, confirming the effectiveness of our approach in expanding the range of feasible manipulation tasks that can be learned using handheld gripper systems, without compromising the cross-embodiment advantages inherent to such systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18757
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MV-UMI: A Scalable Multi-View Interface for Cross-Embodiment Learning
Rayyan, Omar
Abanes, John
Hafez, Mahmoud
Tzes, Anthony
Abu-Dakka, Fares
Robotics
Artificial Intelligence
Recent advances in imitation learning have shown great promise for developing robust robot manipulation policies from demonstrations. However, this promise is contingent on the availability of diverse, high-quality datasets, which are not only challenging and costly to collect but are often constrained to a specific robot embodiment. Portable handheld grippers have recently emerged as intuitive and scalable alternatives to traditional robotic teleoperation methods for data collection. However, their reliance solely on first-person view wrist-mounted cameras often creates limitations in capturing sufficient scene contexts. In this paper, we present MV-UMI (Multi-View Universal Manipulation Interface), a framework that integrates a third-person perspective with the egocentric camera to overcome this limitation. This integration mitigates domain shifts between human demonstration and robot deployment, preserving the cross-embodiment advantages of handheld data-collection devices. Our experimental results, including an ablation study, demonstrate that our MV-UMI framework improves performance in sub-tasks requiring broad scene understanding by approximately 47% across 3 tasks, confirming the effectiveness of our approach in expanding the range of feasible manipulation tasks that can be learned using handheld gripper systems, without compromising the cross-embodiment advantages inherent to such systems.
title MV-UMI: A Scalable Multi-View Interface for Cross-Embodiment Learning
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2509.18757