LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cong, Peishan, Wang, Ziyi, Dou, Zhiyang, Ren, Yiming, Yin, Wei, Cheng, Kai, Sun, Yujing, Long, Xiaoxiao, Zhu, Xinge, Ma, Yuexin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914722233712640
author Cong, Peishan
Wang, Ziyi
Dou, Zhiyang
Ren, Yiming
Yin, Wei
Cheng, Kai
Sun, Yujing
Long, Xiaoxiao
Zhu, Xinge
Ma, Yuexin
author_facet Cong, Peishan
Wang, Ziyi
Dou, Zhiyang
Ren, Yiming
Yin, Wei
Cheng, Kai
Sun, Yujing
Long, Xiaoxiao
Zhu, Xinge
Ma, Yuexin
contents Language-guided scene-aware human motion generation has great significance for entertainment and robotics. In response to the limitations of existing datasets, we introduce LaserHuman, a pioneering dataset engineered to revolutionize Scene-Text-to-Motion research. LaserHuman stands out with its inclusion of genuine human motions within 3D environments, unbounded free-form natural language descriptions, a blend of indoor and outdoor scenarios, and dynamic, ever-changing scenes. Diverse modalities of capture data and rich annotations present great opportunities for the research of conditional motion generation, and can also facilitate the development of real-life applications. Moreover, to generate semantically consistent and physically plausible human motions, we propose a multi-conditional diffusion model, which is simple but effective, achieving state-of-the-art performance on existing datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2403_13307
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment
Cong, Peishan
Wang, Ziyi
Dou, Zhiyang
Ren, Yiming
Yin, Wei
Cheng, Kai
Sun, Yujing
Long, Xiaoxiao
Zhu, Xinge
Ma, Yuexin
Computer Vision and Pattern Recognition
Language-guided scene-aware human motion generation has great significance for entertainment and robotics. In response to the limitations of existing datasets, we introduce LaserHuman, a pioneering dataset engineered to revolutionize Scene-Text-to-Motion research. LaserHuman stands out with its inclusion of genuine human motions within 3D environments, unbounded free-form natural language descriptions, a blend of indoor and outdoor scenarios, and dynamic, ever-changing scenes. Diverse modalities of capture data and rich annotations present great opportunities for the research of conditional motion generation, and can also facilitate the development of real-life applications. Moreover, to generate semantically consistent and physically plausible human motions, we propose a multi-conditional diffusion model, which is simple but effective, achieving state-of-the-art performance on existing datasets.
title LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.13307