Gaussian Flow Bridges for Audio Domain Transfer with Unpaired Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moliner, Eloi, Braun, Sebastian, Gamper, Hannes
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929365799927808
author Moliner, Eloi
Braun, Sebastian
Gamper, Hannes
author_facet Moliner, Eloi
Braun, Sebastian
Gamper, Hannes
contents Audio domain transfer is the process of modifying audio signals to match characteristics of a different domain, while retaining the original content. This paper investigates the potential of Gaussian Flow Bridges, an emerging approach in generative modeling, for this problem. The presented framework addresses the transport problem across different distributions of audio signals through the implementation of a series of two deterministic probability flows. The proposed framework facilitates manipulation of the target distribution properties through a continuous control variable, which defines a certain aspect of the target domain. Notably, this approach does not rely on paired examples for training. To address identified challenges on maintaining the speech content consistent, we recommend a training strategy that incorporates chunk-based minibatch Optimal Transport couplings of data samples and noise. Comparing our unsupervised method with established baselines, we find competitive performance in tasks of reverberation and distortion manipulation. Despite encoutering limitations, the intriguing results obtained in this study underscore potential for further exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19497
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Gaussian Flow Bridges for Audio Domain Transfer with Unpaired Data
Moliner, Eloi
Braun, Sebastian
Gamper, Hannes
Audio and Speech Processing
Machine Learning
Sound
Audio domain transfer is the process of modifying audio signals to match characteristics of a different domain, while retaining the original content. This paper investigates the potential of Gaussian Flow Bridges, an emerging approach in generative modeling, for this problem. The presented framework addresses the transport problem across different distributions of audio signals through the implementation of a series of two deterministic probability flows. The proposed framework facilitates manipulation of the target distribution properties through a continuous control variable, which defines a certain aspect of the target domain. Notably, this approach does not rely on paired examples for training. To address identified challenges on maintaining the speech content consistent, we recommend a training strategy that incorporates chunk-based minibatch Optimal Transport couplings of data samples and noise. Comparing our unsupervised method with established baselines, we find competitive performance in tasks of reverberation and distortion manipulation. Despite encoutering limitations, the intriguing results obtained in this study underscore potential for further exploration.
title Gaussian Flow Bridges for Audio Domain Transfer with Unpaired Data
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2405.19497