SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hong, Seokhyeon, Kim, Chaelin, Yoon, Serin, Nam, Junghyun, Cha, Sihun, Noh, Junyong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909540929241088
author Hong, Seokhyeon
Kim, Chaelin
Yoon, Serin
Nam, Junghyun
Cha, Sihun
Noh, Junyong
author_facet Hong, Seokhyeon
Kim, Chaelin
Yoon, Serin
Nam, Junghyun
Cha, Sihun
Noh, Junyong
contents Text-driven motion generation has advanced significantly with the rise of denoising diffusion models. However, previous methods often oversimplify representations for the skeletal joints, temporal frames, and textual words, limiting their ability to fully capture the information within each modality and their interactions. Moreover, when using pre-trained models for downstream tasks, such as editing, they typically require additional efforts, including manual interventions, optimization, or fine-tuning. In this paper, we introduce a skeleton-aware latent diffusion (SALAD), a model that explicitly captures the intricate inter-relationships between joints, frames, and words. Furthermore, by leveraging cross-attention maps produced during the generation process, we enable attention-based zero-shot text-driven motion editing using a pre-trained SALAD model, requiring no additional user input beyond text prompts. Our approach significantly outperforms previous methods in terms of text-motion alignment without compromising generation quality, and demonstrates practical versatility by providing diverse editing capabilities beyond generation. Code is available at project page.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13836
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing
Hong, Seokhyeon
Kim, Chaelin
Yoon, Serin
Nam, Junghyun
Cha, Sihun
Noh, Junyong
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Text-driven motion generation has advanced significantly with the rise of denoising diffusion models. However, previous methods often oversimplify representations for the skeletal joints, temporal frames, and textual words, limiting their ability to fully capture the information within each modality and their interactions. Moreover, when using pre-trained models for downstream tasks, such as editing, they typically require additional efforts, including manual interventions, optimization, or fine-tuning. In this paper, we introduce a skeleton-aware latent diffusion (SALAD), a model that explicitly captures the intricate inter-relationships between joints, frames, and words. Furthermore, by leveraging cross-attention maps produced during the generation process, we enable attention-based zero-shot text-driven motion editing using a pre-trained SALAD model, requiring no additional user input beyond text prompts. Our approach significantly outperforms previous methods in terms of text-motion alignment without compromising generation quality, and demonstrates practical versatility by providing diverse editing capabilities beyond generation. Code is available at project page.
title SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2503.13836