YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912028176678912 |
|---|---|
| author | Zhou, Xuanru Kashyap, Anshul Li, Steve Sharma, Ayati Morin, Brittany Baquirin, David Vonk, Jet Ezzes, Zoe Miller, Zachary Tempini, Maria Luisa Gorno Lian, Jiachen Anumanchipalli, Gopala Krishna |
| author_facet | Zhou, Xuanru Kashyap, Anshul Li, Steve Sharma, Ayati Morin, Brittany Baquirin, David Vonk, Jet Ezzes, Zoe Miller, Zachary Tempini, Maria Luisa Gorno Lian, Jiachen Anumanchipalli, Gopala Krishna |
| contents | Dysfluent speech detection is the bottleneck for disordered speech analysis and spoken language learning. Current state-of-the-art models are governed by rule-based systems which lack efficiency and robustness, and are sensitive to template design. In this paper, we propose YOLO-Stutter: a first end-to-end method that detects dysfluencies in a time-accurate manner. YOLO-Stutter takes imperfect speech-text alignment as input, followed by a spatial feature aggregator, and a temporal dependency extractor to perform region-wise boundary and class predictions. We also introduce two dysfluency corpus, VCTK-Stutter and VCTK-TTS, that simulate natural spoken dysfluencies including repetition, block, missing, replacement, and prolongation. Our end-to-end method achieves state-of-the-art performance with a minimum number of trainable parameters for on both simulated data and real aphasia speech. Code and datasets are open-sourced at https://github.com/rorizzz/YOLO-Stutter |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_15297 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection Zhou, Xuanru Kashyap, Anshul Li, Steve Sharma, Ayati Morin, Brittany Baquirin, David Vonk, Jet Ezzes, Zoe Miller, Zachary Tempini, Maria Luisa Gorno Lian, Jiachen Anumanchipalli, Gopala Krishna Audio and Speech Processing Artificial Intelligence Computation and Language Dysfluent speech detection is the bottleneck for disordered speech analysis and spoken language learning. Current state-of-the-art models are governed by rule-based systems which lack efficiency and robustness, and are sensitive to template design. In this paper, we propose YOLO-Stutter: a first end-to-end method that detects dysfluencies in a time-accurate manner. YOLO-Stutter takes imperfect speech-text alignment as input, followed by a spatial feature aggregator, and a temporal dependency extractor to perform region-wise boundary and class predictions. We also introduce two dysfluency corpus, VCTK-Stutter and VCTK-TTS, that simulate natural spoken dysfluencies including repetition, block, missing, replacement, and prolongation. Our end-to-end method achieves state-of-the-art performance with a minimum number of trainable parameters for on both simulated data and real aphasia speech. Code and datasets are open-sourced at https://github.com/rorizzz/YOLO-Stutter |
| title | YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection |
| topic | Audio and Speech Processing Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2408.15297 |