EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Jianrong, Fan, Hehe, Yang, Yi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912412386459648
author Zhang, Jianrong
Fan, Hehe
Yang, Yi
author_facet Zhang, Jianrong
Fan, Hehe
Yang, Yi
contents Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diffusion models to effectively compose multiple semantic concepts into a single, coherent motion sequence. To address this issue, we propose EnergyMoGen, which includes two spectrums of Energy-Based Models: (1) We interpret the diffusion model as a latent-aware energy-based model that generates motions by composing a set of diffusion models in latent space; (2) We introduce a semantic-aware energy model based on cross-attention, which enables semantic composition and adaptive gradient descent for text embeddings. To overcome the challenges of semantic inconsistency and motion distortion across these two spectrums, we introduce Synergistic Energy Fusion. This design allows the motion latent diffusion model to synthesize high-quality, complex motions by combining multiple energy terms corresponding to textual descriptions. Experiments show that our approach outperforms existing state-of-the-art models on various motion generation tasks, including text-to-motion generation, compositional motion generation, and multi-concept motion generation. Additionally, we demonstrate that our method can be used to extend motion datasets and improve the text-to-motion task.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14706
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space
Zhang, Jianrong
Fan, Hehe
Yang, Yi
Computer Vision and Pattern Recognition
Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diffusion models to effectively compose multiple semantic concepts into a single, coherent motion sequence. To address this issue, we propose EnergyMoGen, which includes two spectrums of Energy-Based Models: (1) We interpret the diffusion model as a latent-aware energy-based model that generates motions by composing a set of diffusion models in latent space; (2) We introduce a semantic-aware energy model based on cross-attention, which enables semantic composition and adaptive gradient descent for text embeddings. To overcome the challenges of semantic inconsistency and motion distortion across these two spectrums, we introduce Synergistic Energy Fusion. This design allows the motion latent diffusion model to synthesize high-quality, complex motions by combining multiple energy terms corresponding to textual descriptions. Experiments show that our approach outperforms existing state-of-the-art models on various motion generation tasks, including text-to-motion generation, compositional motion generation, and multi-concept motion generation. Additionally, we demonstrate that our method can be used to extend motion datasets and improve the text-to-motion task.
title EnergyMoGen: Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.14706