Music to Dance as Language Translation using Sequence Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Correia, André, Alexandre, Luís A.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912074851942400
author Correia, André
Alexandre, Luís A.
author_facet Correia, André
Alexandre, Luís A.
contents Synthesising appropriate choreographies from music remains an open problem. We introduce MDLT, a novel approach that frames the choreography generation problem as a translation task. Our method leverages an existing data set to learn to translate sequences of audio into corresponding dance poses. We present two variants of MDLT: one utilising the Transformer architecture and the other employing the Mamba architecture. We train our method on AIST++ and PhantomDance data sets to teach a robotic arm to dance, but our method can be applied to a full humanoid robot. Evaluation metrics, including Average Joint Error and Fréchet Inception Distance, consistently demonstrate that, when given a piece of music, MDLT excels at producing realistic and high-quality choreography. The code can be found at github.com/meowatthemoon/MDLT.
format Preprint
id arxiv_https___arxiv_org_abs_2403_15569
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Music to Dance as Language Translation using Sequence Models
Correia, André
Alexandre, Luís A.
Sound
Robotics
Audio and Speech Processing
Synthesising appropriate choreographies from music remains an open problem. We introduce MDLT, a novel approach that frames the choreography generation problem as a translation task. Our method leverages an existing data set to learn to translate sequences of audio into corresponding dance poses. We present two variants of MDLT: one utilising the Transformer architecture and the other employing the Mamba architecture. We train our method on AIST++ and PhantomDance data sets to teach a robotic arm to dance, but our method can be applied to a full humanoid robot. Evaluation metrics, including Average Joint Error and Fréchet Inception Distance, consistently demonstrate that, when given a piece of music, MDLT excels at producing realistic and high-quality choreography. The code can be found at github.com/meowatthemoon/MDLT.
title Music to Dance as Language Translation using Sequence Models
topic Sound
Robotics
Audio and Speech Processing
url https://arxiv.org/abs/2403.15569