Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916446262525952 |
|---|---|
| author | Yu, Chin-Yun Yeh, Sung-Lin Fazekas, György Tang, Hao |
| author_facet | Yu, Chin-Yun Yeh, Sung-Lin Fazekas, György Tang, Hao |
| contents | Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given low-resolution speech utterances. This is commonly achieved by conditioning the network of noise predictor with low-resolution audio. In this paper, we propose a novel sampling algorithm that communicates the information of the low-resolution audio via the reverse sampling process of DMs. The proposed method can be a drop-in replacement for the vanilla sampling process and can significantly improve the performance of the existing works. Moreover, by coupling the proposed sampling method with an unconditional DM, i.e., a DM with no auxiliary inputs to its noise predictor, we can generalize it to a wide range of SR setups. We also attain state-of-the-art results on the VCTK Multi-Speaker benchmark with this novel formulation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2210_15793 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution Yu, Chin-Yun Yeh, Sung-Lin Fazekas, György Tang, Hao Audio and Speech Processing Sound Signal Processing Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given low-resolution speech utterances. This is commonly achieved by conditioning the network of noise predictor with low-resolution audio. In this paper, we propose a novel sampling algorithm that communicates the information of the low-resolution audio via the reverse sampling process of DMs. The proposed method can be a drop-in replacement for the vanilla sampling process and can significantly improve the performance of the existing works. Moreover, by coupling the proposed sampling method with an unconditional DM, i.e., a DM with no auxiliary inputs to its noise predictor, we can generalize it to a wide range of SR setups. We also attain state-of-the-art results on the VCTK Multi-Speaker benchmark with this novel formulation. |
| title | Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution |
| topic | Audio and Speech Processing Sound Signal Processing |
| url | https://arxiv.org/abs/2210.15793 |