Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Chin-Yun, Yeh, Sung-Lin, Fazekas, György, Tang, Hao
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916446262525952
author Yu, Chin-Yun
Yeh, Sung-Lin
Fazekas, György
Tang, Hao
author_facet Yu, Chin-Yun
Yeh, Sung-Lin
Fazekas, György
Tang, Hao
contents Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given low-resolution speech utterances. This is commonly achieved by conditioning the network of noise predictor with low-resolution audio. In this paper, we propose a novel sampling algorithm that communicates the information of the low-resolution audio via the reverse sampling process of DMs. The proposed method can be a drop-in replacement for the vanilla sampling process and can significantly improve the performance of the existing works. Moreover, by coupling the proposed sampling method with an unconditional DM, i.e., a DM with no auxiliary inputs to its noise predictor, we can generalize it to a wide range of SR setups. We also attain state-of-the-art results on the VCTK Multi-Speaker benchmark with this novel formulation.
format Preprint
id arxiv_https___arxiv_org_abs_2210_15793
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
Yu, Chin-Yun
Yeh, Sung-Lin
Fazekas, György
Tang, Hao
Audio and Speech Processing
Sound
Signal Processing
Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given low-resolution speech utterances. This is commonly achieved by conditioning the network of noise predictor with low-resolution audio. In this paper, we propose a novel sampling algorithm that communicates the information of the low-resolution audio via the reverse sampling process of DMs. The proposed method can be a drop-in replacement for the vanilla sampling process and can significantly improve the performance of the existing works. Moreover, by coupling the proposed sampling method with an unconditional DM, i.e., a DM with no auxiliary inputs to its noise predictor, we can generalize it to a wide range of SR setups. We also attain state-of-the-art results on the VCTK Multi-Speaker benchmark with this novel formulation.
title Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
topic Audio and Speech Processing
Sound
Signal Processing
url https://arxiv.org/abs/2210.15793