A Study on Incorporating Whisper for Robust Speech Assessment

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zezario, Ryandhimas E., Chen, Yu-Wen, Fu, Szu-Wei, Tsao, Yu, Wang, Hsin-Min, Fuh, Chiou-Shann
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913332286455808
author Zezario, Ryandhimas E.
Chen, Yu-Wen
Fu, Szu-Wei
Tsao, Yu
Wang, Hsin-Min
Fuh, Chiou-Shann
author_facet Zezario, Ryandhimas E.
Chen, Yu-Wen
Fu, Szu-Wei
Tsao, Yu
Wang, Hsin-Min
Fuh, Chiou-Shann
contents This research introduces an enhanced version of the multi-objective speech assessment model--MOSA-Net+, by leveraging the acoustic features from Whisper, a large-scaled weakly supervised model. We first investigate the effectiveness of Whisper in deploying a more robust speech assessment model. After that, we explore combining representations from Whisper and SSL models. The experimental results reveal that Whisper's embedding features can contribute to more accurate prediction performance. Moreover, combining the embedding features from Whisper and SSL models only leads to marginal improvement. As compared to intrusive methods, MOSA-Net, and other SSL-based speech assessment models, MOSA-Net+ yields notable improvements in estimating subjective quality and intelligibility scores across all evaluation metrics in Taiwan Mandarin Hearing In Noise test - Quality & Intelligibility (TMHINT-QI) dataset. To further validate its robustness, MOSA-Net+ was tested in the noisy-and-enhanced track of the VoiceMOS Challenge 2023, where it obtained the top-ranked performance among nine systems.
format Preprint
id arxiv_https___arxiv_org_abs_2309_12766
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Study on Incorporating Whisper for Robust Speech Assessment
Zezario, Ryandhimas E.
Chen, Yu-Wen
Fu, Szu-Wei
Tsao, Yu
Wang, Hsin-Min
Fuh, Chiou-Shann
Audio and Speech Processing
Sound
This research introduces an enhanced version of the multi-objective speech assessment model--MOSA-Net+, by leveraging the acoustic features from Whisper, a large-scaled weakly supervised model. We first investigate the effectiveness of Whisper in deploying a more robust speech assessment model. After that, we explore combining representations from Whisper and SSL models. The experimental results reveal that Whisper's embedding features can contribute to more accurate prediction performance. Moreover, combining the embedding features from Whisper and SSL models only leads to marginal improvement. As compared to intrusive methods, MOSA-Net, and other SSL-based speech assessment models, MOSA-Net+ yields notable improvements in estimating subjective quality and intelligibility scores across all evaluation metrics in Taiwan Mandarin Hearing In Noise test - Quality & Intelligibility (TMHINT-QI) dataset. To further validate its robustness, MOSA-Net+ was tested in the noisy-and-enhanced track of the VoiceMOS Challenge 2023, where it obtained the top-ranked performance among nine systems.
title A Study on Incorporating Whisper for Robust Speech Assessment
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2309.12766