Guardado en:
Detalles Bibliográficos
Autores principales: Dang, Shaoxiang, Matsumoto, Tetsuya, Takeuchi, Yoshinori, Tsuboi, Takashi, Tanaka, Yasuhiro, Nakatsubo, Daisuke, Maesawa, Satoshi, Saito, Ryuta, Katsuno, Masahisa, Kudo, Hiroaki
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:https://arxiv.org/abs/2408.12279
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911999479250944
author Dang, Shaoxiang
Matsumoto, Tetsuya
Takeuchi, Yoshinori
Tsuboi, Takashi
Tanaka, Yasuhiro
Nakatsubo, Daisuke
Maesawa, Satoshi
Saito, Ryuta
Katsuno, Masahisa
Kudo, Hiroaki
author_facet Dang, Shaoxiang
Matsumoto, Tetsuya
Takeuchi, Yoshinori
Tsuboi, Takashi
Tanaka, Yasuhiro
Nakatsubo, Daisuke
Maesawa, Satoshi
Saito, Ryuta
Katsuno, Masahisa
Kudo, Hiroaki
contents The potential of deep learning in clinical speech processing is immense, yet the hurdles of limited and imbalanced clinical data samples loom large. This article addresses these challenges by showcasing the utilization of automatic speech recognition and self-supervised learning representations, pre-trained on extensive datasets of normal speech. This innovative approach aims to estimate voice quality of patients with impaired vocal systems. Experiments involve checks on PVQD dataset, covering various causes of vocal system damage in English, and a Japanese dataset focusing on patients with Parkinson's disease before and after undergoing subthalamic nucleus deep brain stimulation (STN-DBS) surgery. The results on PVQD reveal a notable correlation (>0.8 on PCC) and an extraordinary accuracy (<0.5 on MSE) in predicting Grade, Breathy, and Asthenic indicators. Meanwhile, progress has been achieved in predicting the voice quality of patients in the context of STN-DBS.
format Preprint
id arxiv_https___arxiv_org_abs_2408_12279
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Developing vocal system impaired patient-aimed voice quality assessment approach using ASR representation-included multiple features
Dang, Shaoxiang
Matsumoto, Tetsuya
Takeuchi, Yoshinori
Tsuboi, Takashi
Tanaka, Yasuhiro
Nakatsubo, Daisuke
Maesawa, Satoshi
Saito, Ryuta
Katsuno, Masahisa
Kudo, Hiroaki
Sound
Artificial Intelligence
Audio and Speech Processing
The potential of deep learning in clinical speech processing is immense, yet the hurdles of limited and imbalanced clinical data samples loom large. This article addresses these challenges by showcasing the utilization of automatic speech recognition and self-supervised learning representations, pre-trained on extensive datasets of normal speech. This innovative approach aims to estimate voice quality of patients with impaired vocal systems. Experiments involve checks on PVQD dataset, covering various causes of vocal system damage in English, and a Japanese dataset focusing on patients with Parkinson's disease before and after undergoing subthalamic nucleus deep brain stimulation (STN-DBS) surgery. The results on PVQD reveal a notable correlation (>0.8 on PCC) and an extraordinary accuracy (<0.5 on MSE) in predicting Grade, Breathy, and Asthenic indicators. Meanwhile, progress has been achieved in predicting the voice quality of patients in the context of STN-DBS.
title Developing vocal system impaired patient-aimed voice quality assessment approach using ASR representation-included multiple features
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2408.12279