Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lau, Hok-Shing, Huntly, Mark, Morgan, Nathon, Iyenoma, Adesua, Zeng, Biao, Bashford, Tim
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909234814255104
author Lau, Hok-Shing
Huntly, Mark
Morgan, Nathon
Iyenoma, Adesua
Zeng, Biao
Bashford, Tim
author_facet Lau, Hok-Shing
Huntly, Mark
Morgan, Nathon
Iyenoma, Adesua
Zeng, Biao
Bashford, Tim
contents Speech contains information that is clinically relevant to some diseases, which has the potential to be used for health assessment. Recent work shows an interest in applying deep learning algorithms, especially pretrained large speech models to the applications of Automatic Speech Assessment. One question that has not been explored is how these models output the results based on their inputs. In this work, we train and compare two configurations of Audio Spectrogram Transformer in the context of Voice Disorder Detection and apply the attention rollout method to produce model relevance maps, the computed relevance of the spectrogram regions when the model makes predictions. We use these maps to analyse how models make predictions in different conditions and to show that the spread of attention is reduced as a model is finetuned, and the model attention is concentrated on specific phoneme regions.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00531
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
Lau, Hok-Shing
Huntly, Mark
Morgan, Nathon
Iyenoma, Adesua
Zeng, Biao
Bashford, Tim
Sound
Artificial Intelligence
Audio and Speech Processing
Speech contains information that is clinically relevant to some diseases, which has the potential to be used for health assessment. Recent work shows an interest in applying deep learning algorithms, especially pretrained large speech models to the applications of Automatic Speech Assessment. One question that has not been explored is how these models output the results based on their inputs. In this work, we train and compare two configurations of Audio Spectrogram Transformer in the context of Voice Disorder Detection and apply the attention rollout method to produce model relevance maps, the computed relevance of the spectrogram regions when the model makes predictions. We use these maps to analyse how models make predictions in different conditions and to show that the spread of attention is reduced as a model is finetuned, and the model attention is concentrated on specific phoneme regions.
title Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2407.00531