Two-in-One: Unified Multi-Person Interactive Motion Generation by Latent Diffusion Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Boyuan, Wang, Xihua, Song, Ruihua, Huang, Wenbing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916538196426752
author Li, Boyuan
Wang, Xihua
Song, Ruihua
Huang, Wenbing
author_facet Li, Boyuan
Wang, Xihua
Song, Ruihua
Huang, Wenbing
contents Multi-person interactive motion generation, a critical yet under-explored domain in computer character animation, poses significant challenges such as intricate modeling of inter-human interactions beyond individual motions and generating two motions with huge differences from one text condition. Current research often employs separate module branches for individual motions, leading to a loss of interaction information and increased computational demands. To address these challenges, we propose a novel, unified approach that models multi-person motions and their interactions within a single latent space. Our approach streamlines the process by treating interactive motions as an integrated data point, utilizing a Variational AutoEncoder (VAE) for compression into a unified latent space, and performing a diffusion process within this space, guided by the natural language conditions. Experimental results demonstrate our method's superiority over existing approaches in generation quality, performing text condition in particular when motions have significant asymmetry, and accelerating the generation efficiency while preserving high quality.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16670
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Two-in-One: Unified Multi-Person Interactive Motion Generation by Latent Diffusion Transformer
Li, Boyuan
Wang, Xihua
Song, Ruihua
Huang, Wenbing
Computer Vision and Pattern Recognition
Graphics
Multi-person interactive motion generation, a critical yet under-explored domain in computer character animation, poses significant challenges such as intricate modeling of inter-human interactions beyond individual motions and generating two motions with huge differences from one text condition. Current research often employs separate module branches for individual motions, leading to a loss of interaction information and increased computational demands. To address these challenges, we propose a novel, unified approach that models multi-person motions and their interactions within a single latent space. Our approach streamlines the process by treating interactive motions as an integrated data point, utilizing a Variational AutoEncoder (VAE) for compression into a unified latent space, and performing a diffusion process within this space, guided by the natural language conditions. Experimental results demonstrate our method's superiority over existing approaches in generation quality, performing text condition in particular when motions have significant asymmetry, and accelerating the generation efficiency while preserving high quality.
title Two-in-One: Unified Multi-Person Interactive Motion Generation by Latent Diffusion Transformer
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2412.16670