Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913857285390336 |
|---|---|
| author | Guo, Chenxu Lian, Jiachen Zhou, Xuanru Zhang, Jinming Li, Shuhe Ye, Zongli Park, Hwi Joo Das, Anaisha Ezzes, Zoe Vonk, Jet Morin, Brittany Bogley, Rian Wauters, Lisa Miller, Zachary Gorno-Tempini, Maria Anumanchipalli, Gopala |
| author_facet | Guo, Chenxu Lian, Jiachen Zhou, Xuanru Zhang, Jinming Li, Shuhe Ye, Zongli Park, Hwi Joo Das, Anaisha Ezzes, Zoe Vonk, Jet Morin, Brittany Bogley, Rian Wauters, Lisa Miller, Zachary Gorno-Tempini, Maria Anumanchipalli, Gopala |
| contents | Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide insufficient clinical insight, and text-independent models misclassify dysfluency, especially in context-dependent cases. This work introduces Dysfluent-WFST, a zero-shot decoder that simultaneously transcribes phonemes and detects dysfluency. Unlike previous models, Dysfluent-WFST operates with upstream encoders like WavLM and requires no additional training. It achieves state-of-the-art performance in both phonetic error rate and dysfluency detection on simulated and real speech data. Our approach is lightweight, interpretable, and effective, demonstrating that explicit modeling of pronunciation behavior in decoding, rather than complex architectures, is key to improving dysfluency processing systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_16351 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection Guo, Chenxu Lian, Jiachen Zhou, Xuanru Zhang, Jinming Li, Shuhe Ye, Zongli Park, Hwi Joo Das, Anaisha Ezzes, Zoe Vonk, Jet Morin, Brittany Bogley, Rian Wauters, Lisa Miller, Zachary Gorno-Tempini, Maria Anumanchipalli, Gopala Audio and Speech Processing Artificial Intelligence Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide insufficient clinical insight, and text-independent models misclassify dysfluency, especially in context-dependent cases. This work introduces Dysfluent-WFST, a zero-shot decoder that simultaneously transcribes phonemes and detects dysfluency. Unlike previous models, Dysfluent-WFST operates with upstream encoders like WavLM and requires no additional training. It achieves state-of-the-art performance in both phonetic error rate and dysfluency detection on simulated and real speech data. Our approach is lightweight, interpretable, and effective, demonstrating that explicit modeling of pronunciation behavior in decoding, rather than complex architectures, is key to improving dysfluency processing systems. |
| title | Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection |
| topic | Audio and Speech Processing Artificial Intelligence |
| url | https://arxiv.org/abs/2505.16351 |