Towards Open Domain Text-Driven Synthesis of Multi-Person Motions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916324139073536 |
|---|---|
| author | Shan, Mengyi Dong, Lu Han, Yutao Yao, Yuan Liu, Tao Nwogu, Ifeoma Qi, Guo-Jun Hill, Mitch |
| author_facet | Shan, Mengyi Dong, Lu Han, Yutao Yao, Yuan Liu, Tao Nwogu, Ifeoma Qi, Guo-Jun Hill, Mitch |
| contents | This work aims to generate natural and diverse group motions of multiple humans from textual descriptions. While single-person text-to-motion generation is extensively studied, it remains challenging to synthesize motions for more than one or two subjects from in-the-wild prompts, mainly due to the lack of available datasets. In this work, we curate human pose and motion datasets by estimating pose information from large-scale image and video datasets. Our models use a transformer-based diffusion framework that accommodates multiple datasets with any number of subjects or frames. Experiments explore both generation of multi-person static poses and generation of multi-person motion sequences. To our knowledge, our method is the first to generate multi-subject motion sequences with high diversity and fidelity from a large variety of textual prompts. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_18483 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Towards Open Domain Text-Driven Synthesis of Multi-Person Motions Shan, Mengyi Dong, Lu Han, Yutao Yao, Yuan Liu, Tao Nwogu, Ifeoma Qi, Guo-Jun Hill, Mitch Computer Vision and Pattern Recognition This work aims to generate natural and diverse group motions of multiple humans from textual descriptions. While single-person text-to-motion generation is extensively studied, it remains challenging to synthesize motions for more than one or two subjects from in-the-wild prompts, mainly due to the lack of available datasets. In this work, we curate human pose and motion datasets by estimating pose information from large-scale image and video datasets. Our models use a transformer-based diffusion framework that accommodates multiple datasets with any number of subjects or frames. Experiments explore both generation of multi-person static poses and generation of multi-person motion sequences. To our knowledge, our method is the first to generate multi-subject motion sequences with high diversity and fidelity from a large variety of textual prompts. |
| title | Towards Open Domain Text-Driven Synthesis of Multi-Person Motions |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2405.18483 |