HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lv, Xintao, Xu, Liang, Yan, Yichao, Jin, Xin, Xu, Congsheng, Wu, Shuwen, Liu, Yifan, Li, Lincheng, Bi, Mengxiao, Zeng, Wenjun, Yang, Xiaokang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914947055747072
author Lv, Xintao
Xu, Liang
Yan, Yichao
Jin, Xin
Xu, Congsheng
Wu, Shuwen
Liu, Yifan
Li, Lincheng
Bi, Mengxiao
Zeng, Wenjun
Yang, Xiaokang
author_facet Lv, Xintao
Xu, Liang
Yan, Yichao
Jin, Xin
Xu, Congsheng
Wu, Shuwen
Liu, Yifan
Li, Lincheng
Bi, Mengxiao
Zeng, Wenjun
Yang, Xiaokang
contents Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of multiple objects. Thus, we propose HIMO, a large-scale MoCap dataset of full-body human interacting with multiple objects, containing 3.3K 4D HOI sequences and 4.08M 3D HOI frames. We also annotate HIMO with detailed textual descriptions and temporal segments, benchmarking two novel tasks of HOI synthesis conditioned on either the whole text prompt or the segmented text prompts as fine-grained timeline control. To address these novel tasks, we propose a dual-branch conditional diffusion model with a mutual interaction module for HOI synthesis. Besides, an auto-regressive generation pipeline is also designed to obtain smooth transitions between HOI segments. Experimental results demonstrate the generalization ability to unseen object geometries and temporal compositions.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12371
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects
Lv, Xintao
Xu, Liang
Yan, Yichao
Jin, Xin
Xu, Congsheng
Wu, Shuwen
Liu, Yifan
Li, Lincheng
Bi, Mengxiao
Zeng, Wenjun
Yang, Xiaokang
Computer Vision and Pattern Recognition
Artificial Intelligence
Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of multiple objects. Thus, we propose HIMO, a large-scale MoCap dataset of full-body human interacting with multiple objects, containing 3.3K 4D HOI sequences and 4.08M 3D HOI frames. We also annotate HIMO with detailed textual descriptions and temporal segments, benchmarking two novel tasks of HOI synthesis conditioned on either the whole text prompt or the segmented text prompts as fine-grained timeline control. To address these novel tasks, we propose a dual-branch conditional diffusion model with a mutual interaction module for HOI synthesis. Besides, an auto-regressive generation pipeline is also designed to obtain smooth transitions between HOI segments. Experimental results demonstrate the generalization ability to unseen object geometries and temporal compositions.
title HIMO: A New Benchmark for Full-Body Human Interacting with Multiple Objects
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2407.12371