Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908415583846400 |
|---|---|
| author | Zhang, Jinming Zhou, Xuanru Lian, Jiachen Li, Shuhe Li, William Ezzes, Zoe Bogley, Rian Wauters, Lisa Miller, Zachary Vonk, Jet Morin, Brittany Gorno-Tempini, Maria Anumanchipalli, Gopala |
| author_facet | Zhang, Jinming Zhou, Xuanru Lian, Jiachen Li, Shuhe Li, William Ezzes, Zoe Bogley, Rian Wauters, Lisa Miller, Zachary Vonk, Jet Morin, Brittany Gorno-Tempini, Maria Anumanchipalli, Gopala |
| contents | Speech dysfluency detection is crucial for clinical diagnosis and language assessment, but existing methods are limited by the scarcity of high-quality annotated data. Although recent advances in TTS model have enabled synthetic dysfluency generation, existing synthetic datasets suffer from unnatural prosody and limited contextual diversity. To address these limitations, we propose LLM-Dys -- the most comprehensive dysfluent speech corpus with LLM-enhanced dysfluency simulation. This dataset captures 11 dysfluency categories spanning both word and phoneme levels. Building upon this resource, we improve an end-to-end dysfluency detection framework. Experimental validation demonstrates state-of-the-art performance. All data, models, and code are open-sourced at https://github.com/Berkeley-Speech-Group/LLM-Dys. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_22029 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection Zhang, Jinming Zhou, Xuanru Lian, Jiachen Li, Shuhe Li, William Ezzes, Zoe Bogley, Rian Wauters, Lisa Miller, Zachary Vonk, Jet Morin, Brittany Gorno-Tempini, Maria Anumanchipalli, Gopala Audio and Speech Processing Artificial Intelligence Sound Speech dysfluency detection is crucial for clinical diagnosis and language assessment, but existing methods are limited by the scarcity of high-quality annotated data. Although recent advances in TTS model have enabled synthetic dysfluency generation, existing synthetic datasets suffer from unnatural prosody and limited contextual diversity. To address these limitations, we propose LLM-Dys -- the most comprehensive dysfluent speech corpus with LLM-enhanced dysfluency simulation. This dataset captures 11 dysfluency categories spanning both word and phoneme levels. Building upon this resource, we improve an end-to-end dysfluency detection framework. Experimental validation demonstrates state-of-the-art performance. All data, models, and code are open-sourced at https://github.com/Berkeley-Speech-Group/LLM-Dys. |
| title | Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection |
| topic | Audio and Speech Processing Artificial Intelligence Sound |
| url | https://arxiv.org/abs/2505.22029 |