Not Like Transformers: Drop the Beat Representation for Dance Generation with Mamba-Based Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Sangjune, Choi, Inhyeok, Soon, Donghyeon, Jeon, Youngwoo, Joo, Kyungdon
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908873140469760
author Park, Sangjune
Choi, Inhyeok
Soon, Donghyeon
Jeon, Youngwoo
Joo, Kyungdon
author_facet Park, Sangjune
Choi, Inhyeok
Soon, Donghyeon
Jeon, Youngwoo
Joo, Kyungdon
contents Dance is a form of human motion characterized by emotional expression and communication, playing a role in various fields such as music, virtual reality, and content creation. Existing methods for dance generation often fail to adequately capture the inherently sequential, rhythmical, and music-synchronized characteristics of dance. In this paper, we propose \emph{MambaDance}, a new dance generation approach that leverages a Mamba-based diffusion model. Mamba, well-suited to handling long and autoregressive sequences, is integrated into our two-stage diffusion architecture, substituting off-the-shelf Transformer. Additionally, considering the critical role of musical beats in dance choreography, we propose a Gaussian-based beat representation to explicitly guide the decoding of dance sequences. Experiments on AIST++ and FineDance datasets for each sequence length show that our proposed method effectively generates plausible dance movements while reflecting essential characteristics, consistently from short to long dances, compared to the previous methods. Additional qualitative results and demo videos are available at \small{https://vision3d-lab.github.io/mambadance}.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08023
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Not Like Transformers: Drop the Beat Representation for Dance Generation with Mamba-Based Diffusion Model
Park, Sangjune
Choi, Inhyeok
Soon, Donghyeon
Jeon, Youngwoo
Joo, Kyungdon
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Sound
Dance is a form of human motion characterized by emotional expression and communication, playing a role in various fields such as music, virtual reality, and content creation. Existing methods for dance generation often fail to adequately capture the inherently sequential, rhythmical, and music-synchronized characteristics of dance. In this paper, we propose \emph{MambaDance}, a new dance generation approach that leverages a Mamba-based diffusion model. Mamba, well-suited to handling long and autoregressive sequences, is integrated into our two-stage diffusion architecture, substituting off-the-shelf Transformer. Additionally, considering the critical role of musical beats in dance choreography, we propose a Gaussian-based beat representation to explicitly guide the decoding of dance sequences. Experiments on AIST++ and FineDance datasets for each sequence length show that our proposed method effectively generates plausible dance movements while reflecting essential characteristics, consistently from short to long dances, compared to the previous methods. Additional qualitative results and demo videos are available at \small{https://vision3d-lab.github.io/mambadance}.
title Not Like Transformers: Drop the Beat Representation for Dance Generation with Mamba-Based Diffusion Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Sound
url https://arxiv.org/abs/2603.08023