WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Akinrintoyo, Emmanuel, Abdelhalim, Nadine, Salomons, Nicole
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909625402523648
author Akinrintoyo, Emmanuel
Abdelhalim, Nadine
Salomons, Nicole
author_facet Akinrintoyo, Emmanuel
Abdelhalim, Nadine
Salomons, Nicole
contents Whisper fails to correctly transcribe dementia speech because persons with dementia (PwDs) often exhibit irregular speech patterns and disfluencies such as pauses, repetitions, and fragmented sentences. It was trained on standard speech and may have had little or no exposure to dementia-affected speech. However, correct transcription is vital for dementia speech for cost-effective diagnosis and the development of assistive technology. In this work, we fine-tune Whisper with the open-source dementia speech dataset (DementiaBank) and our in-house dataset to improve its word error rate (WER). The fine-tuning also includes filler words to ascertain the filler inclusion rate (FIR) and F1 score. The fine-tuned models significantly outperformed the off-the-shelf models. The medium-sized model achieved a WER of 0.24, outperforming previous work. Similarly, there was a notable generalisability to unseen data and speech patterns.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21551
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
Akinrintoyo, Emmanuel
Abdelhalim, Nadine
Salomons, Nicole
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
Whisper fails to correctly transcribe dementia speech because persons with dementia (PwDs) often exhibit irregular speech patterns and disfluencies such as pauses, repetitions, and fragmented sentences. It was trained on standard speech and may have had little or no exposure to dementia-affected speech. However, correct transcription is vital for dementia speech for cost-effective diagnosis and the development of assistive technology. In this work, we fine-tune Whisper with the open-source dementia speech dataset (DementiaBank) and our in-house dataset to improve its word error rate (WER). The fine-tuning also includes filler words to ascertain the filler inclusion rate (FIR) and F1 score. The fine-tuned models significantly outperformed the off-the-shelf models. The medium-sized model achieved a WER of 0.24, outperforming previous work. Similarly, there was a notable generalisability to unseen data and speech patterns.
title WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
Sound
url https://arxiv.org/abs/2505.21551