Let There Be Sound: Reconstructing High Quality Speech from Silent Videos

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Ji-Hoon, Kim, Jaehun, Chung, Joon Son
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910287293054976
author Kim, Ji-Hoon
Kim, Jaehun
Chung, Joon Son
author_facet Kim, Ji-Hoon
Kim, Jaehun
Chung, Joon Son
contents The goal of this work is to reconstruct high quality speech from lip motions alone, a task also known as lip-to-speech. A key challenge of lip-to-speech systems is the one-to-many mapping caused by (1) the existence of homophenes and (2) multiple speech variations, resulting in a mispronounced and over-smoothed speech. In this paper, we propose a novel lip-to-speech system that significantly improves the generation quality by alleviating the one-to-many mapping problem from multiple perspectives. Specifically, we incorporate (1) self-supervised speech representations to disambiguate homophenes, and (2) acoustic variance information to model diverse speech styles. Additionally, to better solve the aforementioned problem, we employ a flow based post-net which captures and refines the details of the generated speech. We perform extensive experiments on two datasets, and demonstrate that our method achieves the generation quality close to that of real human utterance, outperforming existing methods in terms of speech naturalness and intelligibility by a large margin. Synthesised samples are available at our demo page: https://mm.kaist.ac.kr/projects/LTBS.
format Preprint
id arxiv_https___arxiv_org_abs_2308_15256
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
Kim, Ji-Hoon
Kim, Jaehun
Chung, Joon Son
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
The goal of this work is to reconstruct high quality speech from lip motions alone, a task also known as lip-to-speech. A key challenge of lip-to-speech systems is the one-to-many mapping caused by (1) the existence of homophenes and (2) multiple speech variations, resulting in a mispronounced and over-smoothed speech. In this paper, we propose a novel lip-to-speech system that significantly improves the generation quality by alleviating the one-to-many mapping problem from multiple perspectives. Specifically, we incorporate (1) self-supervised speech representations to disambiguate homophenes, and (2) acoustic variance information to model diverse speech styles. Additionally, to better solve the aforementioned problem, we employ a flow based post-net which captures and refines the details of the generated speech. We perform extensive experiments on two datasets, and demonstrate that our method achieves the generation quality close to that of real human utterance, outperforming existing methods in terms of speech naturalness and intelligibility by a large margin. Synthesised samples are available at our demo page: https://mm.kaist.ac.kr/projects/LTBS.
title Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
url https://arxiv.org/abs/2308.15256