FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jung, Chaeyoung, Lee, Suyeon, Kim, Ji-Hoon, Chung, Joon Son
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917693049798656
author Jung, Chaeyoung
Lee, Suyeon
Kim, Ji-Hoon
Chung, Joon Son
author_facet Jung, Chaeyoung
Lee, Suyeon
Kim, Ji-Hoon
Chung, Joon Son
contents This work proposes an efficient method to enhance the quality of corrupted speech signals by leveraging both acoustic and visual cues. While existing diffusion-based approaches have demonstrated remarkable quality, their applicability is limited by slow inference speeds and computational complexity. To address this issue, we present FlowAVSE which enhances the inference speed and reduces the number of learnable parameters without degrading the output quality. In particular, we employ a conditional flow matching algorithm that enables the generation of high-quality speech in a single sampling step. Moreover, we increase efficiency by optimizing the underlying U-net architecture of diffusion-based systems. Our experiments demonstrate that FlowAVSE achieves 22 times faster inference speed and reduces the model size by half while maintaining the output quality. The demo page is available at: https://cyongong.github.io/FlowAVSE.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2406_09286
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
Jung, Chaeyoung
Lee, Suyeon
Kim, Ji-Hoon
Chung, Joon Son
Audio and Speech Processing
Sound
This work proposes an efficient method to enhance the quality of corrupted speech signals by leveraging both acoustic and visual cues. While existing diffusion-based approaches have demonstrated remarkable quality, their applicability is limited by slow inference speeds and computational complexity. To address this issue, we present FlowAVSE which enhances the inference speed and reduces the number of learnable parameters without degrading the output quality. In particular, we employ a conditional flow matching algorithm that enables the generation of high-quality speech in a single sampling step. Moreover, we increase efficiency by optimizing the underlying U-net architecture of diffusion-based systems. Our experiments demonstrate that FlowAVSE achieves 22 times faster inference speed and reduces the model size by half while maintaining the output quality. The demo page is available at: https://cyongong.github.io/FlowAVSE.github.io/
title FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2406.09286