HUMOTO: A 4D Dataset of Mocap Human Object Interactions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Jiaxin, Huang, Chun-Hao Paul, Bhattacharya, Uttaran, Huang, Qixing, Zhou, Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909846627942400
author Lu, Jiaxin
Huang, Chun-Hao Paul
Bhattacharya, Uttaran
Huang, Qixing
Zhou, Yi
author_facet Lu, Jiaxin
Huang, Chun-Hao Paul
Bhattacharya, Uttaran
Huang, Qixing
Zhou, Yi
contents We present Human Motions with Objects (HUMOTO), a high-fidelity dataset of human-object interactions for motion generation, computer vision, and robotics applications. Featuring 735 sequences (7,875 seconds at 30 fps), HUMOTO captures interactions with 63 precisely modeled objects and 72 articulated parts. Our innovations include a scene-driven LLM scripting pipeline creating complete, purposeful tasks with natural progression, and a mocap-and-camera recording setup to effectively handle occlusions. Spanning diverse activities from cooking to outdoor picnics, HUMOTO preserves both physical accuracy and logical task flow. Professional artists rigorously clean and verify each sequence, minimizing foot sliding and object penetrations. We also provide benchmarks compared to other datasets. HUMOTO's comprehensive full-body motion and simultaneous multi-object interactions address key data-capturing challenges and provide opportunities to advance realistic human-object interaction modeling across research domains with practical applications in animation, robotics, and embodied AI systems. Project: https://jiaxin-lu.github.io/humoto/ .
format Preprint
id arxiv_https___arxiv_org_abs_2504_10414
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HUMOTO: A 4D Dataset of Mocap Human Object Interactions
Lu, Jiaxin
Huang, Chun-Hao Paul
Bhattacharya, Uttaran
Huang, Qixing
Zhou, Yi
Computer Vision and Pattern Recognition
We present Human Motions with Objects (HUMOTO), a high-fidelity dataset of human-object interactions for motion generation, computer vision, and robotics applications. Featuring 735 sequences (7,875 seconds at 30 fps), HUMOTO captures interactions with 63 precisely modeled objects and 72 articulated parts. Our innovations include a scene-driven LLM scripting pipeline creating complete, purposeful tasks with natural progression, and a mocap-and-camera recording setup to effectively handle occlusions. Spanning diverse activities from cooking to outdoor picnics, HUMOTO preserves both physical accuracy and logical task flow. Professional artists rigorously clean and verify each sequence, minimizing foot sliding and object penetrations. We also provide benchmarks compared to other datasets. HUMOTO's comprehensive full-body motion and simultaneous multi-object interactions address key data-capturing challenges and provide opportunities to advance realistic human-object interaction modeling across research domains with practical applications in animation, robotics, and embodied AI systems. Project: https://jiaxin-lu.github.io/humoto/ .
title HUMOTO: A 4D Dataset of Mocap Human Object Interactions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.10414