Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Engelbracht, Tim, Zurbrügg, René, Wohlrapp, Matteo, Büchner, Martin, Valada, Abhinav, Pollefeys, Marc, Blum, Hermann, Bauer, Zuria
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917412542087168
author Engelbracht, Tim
Zurbrügg, René
Wohlrapp, Matteo
Büchner, Martin
Valada, Abhinav
Pollefeys, Marc
Blum, Hermann
Bauer, Zuria
author_facet Engelbracht, Tim
Zurbrügg, René
Wohlrapp, Matteo
Büchner, Martin
Valada, Abhinav
Pollefeys, Marc
Blum, Hermann
Bauer, Zuria
contents We present a dataset for force-grounded, cross-view articulated manipulation that couples what is seen with what is done and what is felt during real human interaction. The dataset contains 3048 sequences across 381 articulated objects in 38 environments. Each object is operated in four embodiments - (i) human hand, (ii) human hand with a wrist-mounted camera, (iii) handheld UMI gripper, and (iv) a custom Hoi! gripper, where the tool embodiment provides end-effector forces and tactile sensing. Our dataset offers a holistic view of interaction understanding from video, enabling researchers to evaluate how well methods transfer between human and robotic viewpoints, but also investigate underexplored modalities such as interaction forces. The Project Website can be found at https://timengelbracht.github.io/Hoi-Dataset-Website/.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04884
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
Engelbracht, Tim
Zurbrügg, René
Wohlrapp, Matteo
Büchner, Martin
Valada, Abhinav
Pollefeys, Marc
Blum, Hermann
Bauer, Zuria
Robotics
We present a dataset for force-grounded, cross-view articulated manipulation that couples what is seen with what is done and what is felt during real human interaction. The dataset contains 3048 sequences across 381 articulated objects in 38 environments. Each object is operated in four embodiments - (i) human hand, (ii) human hand with a wrist-mounted camera, (iii) handheld UMI gripper, and (iv) a custom Hoi! gripper, where the tool embodiment provides end-effector forces and tactile sensing. Our dataset offers a holistic view of interaction understanding from video, enabling researchers to evaluate how well methods transfer between human and robotic viewpoints, but also investigate underexplored modalities such as interaction forces. The Project Website can be found at https://timengelbracht.github.io/Hoi-Dataset-Website/.
title Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
topic Robotics
url https://arxiv.org/abs/2512.04884