Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909234814255104 |
|---|---|
| author | Lau, Hok-Shing Huntly, Mark Morgan, Nathon Iyenoma, Adesua Zeng, Biao Bashford, Tim |
| author_facet | Lau, Hok-Shing Huntly, Mark Morgan, Nathon Iyenoma, Adesua Zeng, Biao Bashford, Tim |
| contents | Speech contains information that is clinically relevant to some diseases, which has the potential to be used for health assessment. Recent work shows an interest in applying deep learning algorithms, especially pretrained large speech models to the applications of Automatic Speech Assessment. One question that has not been explored is how these models output the results based on their inputs. In this work, we train and compare two configurations of Audio Spectrogram Transformer in the context of Voice Disorder Detection and apply the attention rollout method to produce model relevance maps, the computed relevance of the spectrogram regions when the model makes predictions. We use these maps to analyse how models make predictions in different conditions and to show that the spread of attention is reduced as a model is finetuned, and the model attention is concentrated on specific phoneme regions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_00531 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders Lau, Hok-Shing Huntly, Mark Morgan, Nathon Iyenoma, Adesua Zeng, Biao Bashford, Tim Sound Artificial Intelligence Audio and Speech Processing Speech contains information that is clinically relevant to some diseases, which has the potential to be used for health assessment. Recent work shows an interest in applying deep learning algorithms, especially pretrained large speech models to the applications of Automatic Speech Assessment. One question that has not been explored is how these models output the results based on their inputs. In this work, we train and compare two configurations of Audio Spectrogram Transformer in the context of Voice Disorder Detection and apply the attention rollout method to produce model relevance maps, the computed relevance of the spectrogram regions when the model makes predictions. We use these maps to analyse how models make predictions in different conditions and to show that the spread of attention is reduced as a model is finetuned, and the model attention is concentrated on specific phoneme regions. |
| title | Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders |
| topic | Sound Artificial Intelligence Audio and Speech Processing |
| url | https://arxiv.org/abs/2407.00531 |