Towards Open Domain Text-Driven Synthesis of Multi-Person Motions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shan, Mengyi, Dong, Lu, Han, Yutao, Yao, Yuan, Liu, Tao, Nwogu, Ifeoma, Qi, Guo-Jun, Hill, Mitch
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916324139073536
author Shan, Mengyi
Dong, Lu
Han, Yutao
Yao, Yuan
Liu, Tao
Nwogu, Ifeoma
Qi, Guo-Jun
Hill, Mitch
author_facet Shan, Mengyi
Dong, Lu
Han, Yutao
Yao, Yuan
Liu, Tao
Nwogu, Ifeoma
Qi, Guo-Jun
Hill, Mitch
contents This work aims to generate natural and diverse group motions of multiple humans from textual descriptions. While single-person text-to-motion generation is extensively studied, it remains challenging to synthesize motions for more than one or two subjects from in-the-wild prompts, mainly due to the lack of available datasets. In this work, we curate human pose and motion datasets by estimating pose information from large-scale image and video datasets. Our models use a transformer-based diffusion framework that accommodates multiple datasets with any number of subjects or frames. Experiments explore both generation of multi-person static poses and generation of multi-person motion sequences. To our knowledge, our method is the first to generate multi-subject motion sequences with high diversity and fidelity from a large variety of textual prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18483
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Open Domain Text-Driven Synthesis of Multi-Person Motions
Shan, Mengyi
Dong, Lu
Han, Yutao
Yao, Yuan
Liu, Tao
Nwogu, Ifeoma
Qi, Guo-Jun
Hill, Mitch
Computer Vision and Pattern Recognition
This work aims to generate natural and diverse group motions of multiple humans from textual descriptions. While single-person text-to-motion generation is extensively studied, it remains challenging to synthesize motions for more than one or two subjects from in-the-wild prompts, mainly due to the lack of available datasets. In this work, we curate human pose and motion datasets by estimating pose information from large-scale image and video datasets. Our models use a transformer-based diffusion framework that accommodates multiple datasets with any number of subjects or frames. Experiments explore both generation of multi-person static poses and generation of multi-person motion sequences. To our knowledge, our method is the first to generate multi-subject motion sequences with high diversity and fidelity from a large variety of textual prompts.
title Towards Open Domain Text-Driven Synthesis of Multi-Person Motions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.18483