MusicInfuser: Making Video Diffusion Listen and Dance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hong, Susung, Kemelmacher-Shlizerman, Ira, Curless, Brian, Seitz, Steven M.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917455241150464
author Hong, Susung
Kemelmacher-Shlizerman, Ira
Curless, Brian
Seitz, Steven M.
author_facet Hong, Susung
Kemelmacher-Shlizerman, Ira
Curless, Brian
Seitz, Steven M.
contents We introduce MusicInfuser, an approach that aligns pre-trained text-to-video diffusion models to generate high-quality dance videos synchronized with specified music tracks. Rather than training a multimodal audio-video or audio-motion model from scratch, our method demonstrates how existing video diffusion models can be efficiently adapted to align with musical inputs. We propose a novel layer-wise adaptability criterion based on a guidance-inspired constructive influence function to select adaptable layers, significantly reducing training costs while preserving rich prior knowledge, even with limited, specialized datasets. Experiments show that MusicInfuser effectively bridges the gap between music and video, generating novel and diverse dance movements that respond dynamically to music. Furthermore, our framework generalizes well to unseen music tracks, longer video sequences, and unconventional subjects, outperforming baseline models in consistency and synchronization. All of this is achieved without requiring motion data, with training completed on a single GPU within a day.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14505
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MusicInfuser: Making Video Diffusion Listen and Dance
Hong, Susung
Kemelmacher-Shlizerman, Ira
Curless, Brian
Seitz, Steven M.
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
We introduce MusicInfuser, an approach that aligns pre-trained text-to-video diffusion models to generate high-quality dance videos synchronized with specified music tracks. Rather than training a multimodal audio-video or audio-motion model from scratch, our method demonstrates how existing video diffusion models can be efficiently adapted to align with musical inputs. We propose a novel layer-wise adaptability criterion based on a guidance-inspired constructive influence function to select adaptable layers, significantly reducing training costs while preserving rich prior knowledge, even with limited, specialized datasets. Experiments show that MusicInfuser effectively bridges the gap between music and video, generating novel and diverse dance movements that respond dynamically to music. Furthermore, our framework generalizes well to unseen music tracks, longer video sequences, and unconventional subjects, outperforming baseline models in consistency and synchronization. All of this is achieved without requiring motion data, with training completed on a single GPU within a day.
title MusicInfuser: Making Video Diffusion Listen and Dance
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.14505