Repurposing Image Diffusion Models for Training-Free Music Style Transfer on Mel-spectrograms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Heehwan, Kwon, Joonwoo, Kim, Sooyoung, Seo, Jungwoo, Yoo, Shinjae, Lin, Yuewei, Cha, Jiook
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916007530987520
author Wang, Heehwan
Kwon, Joonwoo
Kim, Sooyoung
Seo, Jungwoo
Yoo, Shinjae
Lin, Yuewei
Cha, Jiook
author_facet Wang, Heehwan
Kwon, Joonwoo
Kim, Sooyoung
Seo, Jungwoo
Yoo, Shinjae
Lin, Yuewei
Cha, Jiook
contents Music style transfer blends source structure with reference style to enable personalized music creation. However, existing zero-shot methods often struggle to capture fine-grained audio nuances, relying on coarse text descriptions or requiring expensive task-specific training. We propose Stylus, a training-free framework that repurposes pretrained image diffusion models for music style transfer in the Mel-spectrogram domain. By treating audio as structured time-frequency images, Stylus manipulates self-attention by injecting style keys and values while preserving source structural queries. To ensure high fidelity, we introduce a phase-preserving reconstruction strategy to mitigate spectrogram inversion artifacts, alongside a classifier-free-guidance-inspired control for adjustable stylization. Extensive evaluations including 2,925 human ratings demonstrate that Stylus outperforms state-of-the-art baselines, achieving 34.1% higher content preservation and 25.7% better perceptual quality. Our work validates that generic image priors can be effectively leveraged for the training-free transformation of structured Mel-spectrograms. Code and materials are available at https://github.com/Sooyyoungg/Stylus.git.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15913
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Repurposing Image Diffusion Models for Training-Free Music Style Transfer on Mel-spectrograms
Wang, Heehwan
Kwon, Joonwoo
Kim, Sooyoung
Seo, Jungwoo
Yoo, Shinjae
Lin, Yuewei
Cha, Jiook
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Music style transfer blends source structure with reference style to enable personalized music creation. However, existing zero-shot methods often struggle to capture fine-grained audio nuances, relying on coarse text descriptions or requiring expensive task-specific training. We propose Stylus, a training-free framework that repurposes pretrained image diffusion models for music style transfer in the Mel-spectrogram domain. By treating audio as structured time-frequency images, Stylus manipulates self-attention by injecting style keys and values while preserving source structural queries. To ensure high fidelity, we introduce a phase-preserving reconstruction strategy to mitigate spectrogram inversion artifacts, alongside a classifier-free-guidance-inspired control for adjustable stylization. Extensive evaluations including 2,925 human ratings demonstrate that Stylus outperforms state-of-the-art baselines, achieving 34.1% higher content preservation and 25.7% better perceptual quality. Our work validates that generic image priors can be effectively leveraged for the training-free transformation of structured Mel-spectrograms. Code and materials are available at https://github.com/Sooyyoungg/Stylus.git.
title Repurposing Image Diffusion Models for Training-Free Music Style Transfer on Mel-spectrograms
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2411.15913