FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nguyen, Tan Dat, Kim, Ji-Hoon, Jang, Youngjoon, Kim, Jaehun, Chung, Joon Son
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913199539879936
author Nguyen, Tan Dat
Kim, Ji-Hoon
Jang, Youngjoon
Kim, Jaehun
Chung, Joon Son
author_facet Nguyen, Tan Dat
Kim, Ji-Hoon
Jang, Youngjoon
Kim, Jaehun
Chung, Joon Son
contents The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes a complicated waveform into sub-band wavelets, which helps FreGrad to operate on a simple and concise feature space, (2) We design a frequency-aware dilated convolution that elevates frequency awareness, resulting in generating speech with accurate frequency information, and (3) We introduce a bag of tricks that boosts the generation quality of the proposed model. In our experiments, FreGrad achieves 3.7 times faster training time and 2.2 times faster inference speed compared to our baseline while reducing the model size by 0.6 times (only 1.78M parameters) without sacrificing the output quality. Audio samples are available at: https://mm.kaist.ac.kr/projects/FreGrad.
format Preprint
id arxiv_https___arxiv_org_abs_2401_10032
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
Nguyen, Tan Dat
Kim, Ji-Hoon
Jang, Youngjoon
Kim, Jaehun
Chung, Joon Son
Audio and Speech Processing
Artificial Intelligence
Signal Processing
The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes a complicated waveform into sub-band wavelets, which helps FreGrad to operate on a simple and concise feature space, (2) We design a frequency-aware dilated convolution that elevates frequency awareness, resulting in generating speech with accurate frequency information, and (3) We introduce a bag of tricks that boosts the generation quality of the proposed model. In our experiments, FreGrad achieves 3.7 times faster training time and 2.2 times faster inference speed compared to our baseline while reducing the model size by 0.6 times (only 1.78M parameters) without sacrificing the output quality. Audio samples are available at: https://mm.kaist.ac.kr/projects/FreGrad.
title FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
topic Audio and Speech Processing
Artificial Intelligence
Signal Processing
url https://arxiv.org/abs/2401.10032