Huntington Disease Automatic Speech Recognition with Biomarker Supervision

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Charles L., Chen, Cady, Gong, Ziwei, Hirschberg, Julia
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915855065939968
author Wang, Charles L.
Chen, Cady
Gong, Ziwei
Hirschberg, Julia
author_facet Wang, Charles L.
Chen, Cady
Gong, Ziwei
Hirschberg, Julia
contents Automatic speech recognition (ASR) for pathological speech remains underexplored, especially for Huntington's disease (HD), where irregular timing, unstable phonation, and articulatory distortion challenge current models. We present a systematic HD-ASR study using a high-fidelity clinical speech corpus not previously used for end-to-end ASR training. We compare multiple ASR families under a unified evaluation, analyzing WER as well as substitution, deletion, and insertion patterns. HD speech induces architecture-specific error regimes, with Parakeet-TDT outperforming encoder-decoder and CTC baselines. HD-specific adaptation reduces WER from 6.99% to 4.95% and we also propose a method for using biomarker-based auxiliary supervision and analyze how error behavior is reshaped in severity-dependent ways rather than uniformly improving WER. We open-source all code and models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_11168
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Huntington Disease Automatic Speech Recognition with Biomarker Supervision
Wang, Charles L.
Chen, Cady
Gong, Ziwei
Hirschberg, Julia
Machine Learning
Computation and Language
Sound
Automatic speech recognition (ASR) for pathological speech remains underexplored, especially for Huntington's disease (HD), where irregular timing, unstable phonation, and articulatory distortion challenge current models. We present a systematic HD-ASR study using a high-fidelity clinical speech corpus not previously used for end-to-end ASR training. We compare multiple ASR families under a unified evaluation, analyzing WER as well as substitution, deletion, and insertion patterns. HD speech induces architecture-specific error regimes, with Parakeet-TDT outperforming encoder-decoder and CTC baselines. HD-specific adaptation reduces WER from 6.99% to 4.95% and we also propose a method for using biomarker-based auxiliary supervision and analyze how error behavior is reshaped in severity-dependent ways rather than uniformly improving WER. We open-source all code and models.
title Huntington Disease Automatic Speech Recognition with Biomarker Supervision
topic Machine Learning
Computation and Language
Sound
url https://arxiv.org/abs/2603.11168