WaveTransfer: A Flexible End-to-end Multi-instrument Timbre Transfer with Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baoueb, Teysir, Bie, Xiaoyu, Janati, Hicham, Richard, Gael
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913515599560704
author Baoueb, Teysir
Bie, Xiaoyu
Janati, Hicham
Richard, Gael
author_facet Baoueb, Teysir
Bie, Xiaoyu
Janati, Hicham
Richard, Gael
contents As diffusion-based deep generative models gain prevalence, researchers are actively investigating their potential applications across various domains, including music synthesis and style alteration. Within this work, we are interested in timbre transfer, a process that involves seamlessly altering the instrumental characteristics of musical pieces while preserving essential musical elements. This paper introduces WaveTransfer, an end-to-end diffusion model designed for timbre transfer. We specifically employ the bilateral denoising diffusion model (BDDM) for noise scheduling search. Our model is capable of conducting timbre transfer between audio mixtures as well as individual instruments. Notably, it exhibits versatility in that it accommodates multiple types of timbre transfer between unique instrument pairs in a single model, eliminating the need for separate model training for each pairing. Furthermore, unlike recent works limited to 16 kHz, WaveTransfer can be trained at various sampling rates, including the industry-standard 44.1 kHz, a feature of particular interest to the music community.
format Preprint
id arxiv_https___arxiv_org_abs_2409_15321
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle WaveTransfer: A Flexible End-to-end Multi-instrument Timbre Transfer with Diffusion
Baoueb, Teysir
Bie, Xiaoyu
Janati, Hicham
Richard, Gael
Audio and Speech Processing
Sound
As diffusion-based deep generative models gain prevalence, researchers are actively investigating their potential applications across various domains, including music synthesis and style alteration. Within this work, we are interested in timbre transfer, a process that involves seamlessly altering the instrumental characteristics of musical pieces while preserving essential musical elements. This paper introduces WaveTransfer, an end-to-end diffusion model designed for timbre transfer. We specifically employ the bilateral denoising diffusion model (BDDM) for noise scheduling search. Our model is capable of conducting timbre transfer between audio mixtures as well as individual instruments. Notably, it exhibits versatility in that it accommodates multiple types of timbre transfer between unique instrument pairs in a single model, eliminating the need for separate model training for each pairing. Furthermore, unlike recent works limited to 16 kHz, WaveTransfer can be trained at various sampling rates, including the industry-standard 44.1 kHz, a feature of particular interest to the music community.
title WaveTransfer: A Flexible End-to-end Multi-instrument Timbre Transfer with Diffusion
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2409.15321