Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Fu, Szu-Wei, Hung, Kuo-Hsuan, Tsao, Yu, Wang, Yu-Chiang Frank
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929256659943424
author Fu, Szu-Wei
Hung, Kuo-Hsuan
Tsao, Yu
Wang, Yu-Chiang Frank
author_facet Fu, Szu-Wei
Hung, Kuo-Hsuan
Tsao, Yu
Wang, Yu-Chiang Frank
contents Speech quality estimation has recently undergone a paradigm shift from human-hearing expert designs to machine-learning models. However, current models rely mainly on supervised learning, which is time-consuming and expensive for label collection. To solve this problem, we propose VQScore, a self-supervised metric for evaluating speech based on the quantization error of a vector-quantized-variational autoencoder (VQ-VAE). The training of VQ-VAE relies on clean speech; hence, large quantization errors can be expected when the speech is distorted. To further improve correlation with real quality scores, domain knowledge of speech processing is incorporated into the model design. We found that the vector quantization mechanism could also be used for self-supervised speech enhancement (SE) model training. To improve the robustness of the encoder for SE, a novel self-distillation mechanism combined with adversarial training is introduced. In summary, the proposed speech quality estimation method and enhancement models require only clean speech for training without any label requirements. Experimental results show that the proposed VQScore and enhancement model are competitive with supervised baselines. The code will be released after publication.
format Preprint
id arxiv_https___arxiv_org_abs_2402_16321
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
Fu, Szu-Wei
Hung, Kuo-Hsuan
Tsao, Yu
Wang, Yu-Chiang Frank
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Speech quality estimation has recently undergone a paradigm shift from human-hearing expert designs to machine-learning models. However, current models rely mainly on supervised learning, which is time-consuming and expensive for label collection. To solve this problem, we propose VQScore, a self-supervised metric for evaluating speech based on the quantization error of a vector-quantized-variational autoencoder (VQ-VAE). The training of VQ-VAE relies on clean speech; hence, large quantization errors can be expected when the speech is distorted. To further improve correlation with real quality scores, domain knowledge of speech processing is incorporated into the model design. We found that the vector quantization mechanism could also be used for self-supervised speech enhancement (SE) model training. To improve the robustness of the encoder for SE, a novel self-distillation mechanism combined with adversarial training is introduced. In summary, the proposed speech quality estimation method and enhancement models require only clean speech for training without any label requirements. Experimental results show that the proposed VQScore and enhancement model are competitive with supervised baselines. The code will be released after publication.
title Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2402.16321