Tool-as-Interface: Learning Robot Policies from Observing Human Tool Use

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Haonan, Zhu, Cheng, Liu, Shuijing, Li, Yunzhu, Driggs-Campbell, Katherine
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916948479049728
author Chen, Haonan
Zhu, Cheng
Liu, Shuijing
Li, Yunzhu
Driggs-Campbell, Katherine
author_facet Chen, Haonan
Zhu, Cheng
Liu, Shuijing
Li, Yunzhu
Driggs-Campbell, Katherine
contents Tool use is essential for enabling robots to perform complex real-world tasks, but learning such skills requires extensive datasets. While teleoperation is widely used, it is slow, delay-sensitive, and poorly suited for dynamic tasks. In contrast, human videos provide a natural way for data collection without specialized hardware, though they pose challenges on robot learning due to viewpoint variations and embodiment gaps. To address these challenges, we propose a framework that transfers tool-use knowledge from humans to robots. To improve the policy's robustness to viewpoint variations, we use two RGB cameras to reconstruct 3D scenes and apply Gaussian splatting for novel view synthesis. We reduce the embodiment gap using segmented observations and tool-centric, task-space actions to achieve embodiment-invariant visuomotor policy learning. We demonstrate our framework's effectiveness across a diverse suite of tool-use tasks, where our learned policy shows strong generalization and robustness to human perturbations, camera motion, and robot base movement. Our method achieves a 71\% improvement in task success over teleoperation-based diffusion policies and dramatically reduces data collection time by 77\% and 41\% compared to teleoperation and the state-of-the-art interface, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2504_04612
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tool-as-Interface: Learning Robot Policies from Observing Human Tool Use
Chen, Haonan
Zhu, Cheng
Liu, Shuijing
Li, Yunzhu
Driggs-Campbell, Katherine
Robotics
Artificial Intelligence
Machine Learning
Tool use is essential for enabling robots to perform complex real-world tasks, but learning such skills requires extensive datasets. While teleoperation is widely used, it is slow, delay-sensitive, and poorly suited for dynamic tasks. In contrast, human videos provide a natural way for data collection without specialized hardware, though they pose challenges on robot learning due to viewpoint variations and embodiment gaps. To address these challenges, we propose a framework that transfers tool-use knowledge from humans to robots. To improve the policy's robustness to viewpoint variations, we use two RGB cameras to reconstruct 3D scenes and apply Gaussian splatting for novel view synthesis. We reduce the embodiment gap using segmented observations and tool-centric, task-space actions to achieve embodiment-invariant visuomotor policy learning. We demonstrate our framework's effectiveness across a diverse suite of tool-use tasks, where our learned policy shows strong generalization and robustness to human perturbations, camera motion, and robot base movement. Our method achieves a 71\% improvement in task success over teleoperation-based diffusion policies and dramatically reduces data collection time by 77\% and 41\% compared to teleoperation and the state-of-the-art interface, respectively.
title Tool-as-Interface: Learning Robot Policies from Observing Human Tool Use
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.04612