Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Chenxu, Lian, Jiachen, Zhou, Xuanru, Zhang, Jinming, Li, Shuhe, Ye, Zongli, Park, Hwi Joo, Das, Anaisha, Ezzes, Zoe, Vonk, Jet, Morin, Brittany, Bogley, Rian, Wauters, Lisa, Miller, Zachary, Gorno-Tempini, Maria, Anumanchipalli, Gopala
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913857285390336
author Guo, Chenxu
Lian, Jiachen
Zhou, Xuanru
Zhang, Jinming
Li, Shuhe
Ye, Zongli
Park, Hwi Joo
Das, Anaisha
Ezzes, Zoe
Vonk, Jet
Morin, Brittany
Bogley, Rian
Wauters, Lisa
Miller, Zachary
Gorno-Tempini, Maria
Anumanchipalli, Gopala
author_facet Guo, Chenxu
Lian, Jiachen
Zhou, Xuanru
Zhang, Jinming
Li, Shuhe
Ye, Zongli
Park, Hwi Joo
Das, Anaisha
Ezzes, Zoe
Vonk, Jet
Morin, Brittany
Bogley, Rian
Wauters, Lisa
Miller, Zachary
Gorno-Tempini, Maria
Anumanchipalli, Gopala
contents Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide insufficient clinical insight, and text-independent models misclassify dysfluency, especially in context-dependent cases. This work introduces Dysfluent-WFST, a zero-shot decoder that simultaneously transcribes phonemes and detects dysfluency. Unlike previous models, Dysfluent-WFST operates with upstream encoders like WavLM and requires no additional training. It achieves state-of-the-art performance in both phonetic error rate and dysfluency detection on simulated and real speech data. Our approach is lightweight, interpretable, and effective, demonstrating that explicit modeling of pronunciation behavior in decoding, rather than complex architectures, is key to improving dysfluency processing systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16351
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection
Guo, Chenxu
Lian, Jiachen
Zhou, Xuanru
Zhang, Jinming
Li, Shuhe
Ye, Zongli
Park, Hwi Joo
Das, Anaisha
Ezzes, Zoe
Vonk, Jet
Morin, Brittany
Bogley, Rian
Wauters, Lisa
Miller, Zachary
Gorno-Tempini, Maria
Anumanchipalli, Gopala
Audio and Speech Processing
Artificial Intelligence
Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide insufficient clinical insight, and text-independent models misclassify dysfluency, especially in context-dependent cases. This work introduces Dysfluent-WFST, a zero-shot decoder that simultaneously transcribes phonemes and detects dysfluency. Unlike previous models, Dysfluent-WFST operates with upstream encoders like WavLM and requires no additional training. It achieves state-of-the-art performance in both phonetic error rate and dysfluency detection on simulated and real speech data. Our approach is lightweight, interpretable, and effective, demonstrating that explicit modeling of pronunciation behavior in decoding, rather than complex architectures, is key to improving dysfluency processing systems.
title Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection
topic Audio and Speech Processing
Artificial Intelligence
url https://arxiv.org/abs/2505.16351