Probing Whisper for Dysarthric Speech in Detection and Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yue, Zhengjun, Kayande, Devendra, Cvetkovic, Zoran, Loweimi, Erfan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912630231269376
author Yue, Zhengjun
Kayande, Devendra
Cvetkovic, Zoran
Loweimi, Erfan
author_facet Yue, Zhengjun
Kayande, Devendra
Cvetkovic, Zoran
Loweimi, Erfan
contents Large-scale end-to-end models such as Whisper have shown strong performance on diverse speech tasks, but their internal behavior on pathological speech remains poorly understood. Understanding how dysarthric speech is represented across layers is critical for building reliable and explainable clinical assessment tools. This study probes the Whisper-Medium model encoder for dysarthric speech for detection and assessment (i.e., severity classification). We evaluate layer-wise embeddings with a linear classifier under both single-task and multi-task settings, and complement these results with Silhouette scores and mutual information to provide perspectives on layer informativeness. To examine adaptability, we repeat the analysis after fine-tuning Whisper on a dysarthric speech recognition task. Across metrics, the mid-level encoder layers (13-15) emerge as most informative, while fine-tuning induces only modest changes. The findings improve the interpretability of Whisper's embeddings and highlight the potential of probing analyses to guide the use of large-scale pretrained models for pathological speech.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04219
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Probing Whisper for Dysarthric Speech in Detection and Assessment
Yue, Zhengjun
Kayande, Devendra
Cvetkovic, Zoran
Loweimi, Erfan
Audio and Speech Processing
Sound
Large-scale end-to-end models such as Whisper have shown strong performance on diverse speech tasks, but their internal behavior on pathological speech remains poorly understood. Understanding how dysarthric speech is represented across layers is critical for building reliable and explainable clinical assessment tools. This study probes the Whisper-Medium model encoder for dysarthric speech for detection and assessment (i.e., severity classification). We evaluate layer-wise embeddings with a linear classifier under both single-task and multi-task settings, and complement these results with Silhouette scores and mutual information to provide perspectives on layer informativeness. To examine adaptability, we repeat the analysis after fine-tuning Whisper on a dysarthric speech recognition task. Across metrics, the mid-level encoder layers (13-15) emerge as most informative, while fine-tuning induces only modest changes. The findings improve the interpretability of Whisper's embeddings and highlight the potential of probing analyses to guide the use of large-scale pretrained models for pathological speech.
title Probing Whisper for Dysarthric Speech in Detection and Assessment
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2510.04219