Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917775687024640 |
|---|---|
| author | Zhou, Xuanru Cho, Cheol Jun Sharma, Ayati Morin, Brittany Baquirin, David Vonk, Jet Ezzes, Zoe Miller, Zachary Tee, Boon Lead Tempini, Maria Luisa Gorno Lian, Jiachen Anumanchipalli, Gopala |
| author_facet | Zhou, Xuanru Cho, Cheol Jun Sharma, Ayati Morin, Brittany Baquirin, David Vonk, Jet Ezzes, Zoe Miller, Zachary Tee, Boon Lead Tempini, Maria Luisa Gorno Lian, Jiachen Anumanchipalli, Gopala |
| contents | Current de-facto dysfluency modeling methods utilize template matching algorithms which are not generalizable to out-of-domain real-world dysfluencies across languages, and are not scalable with increasing amounts of training data. To handle these problems, we propose Stutter-Solver: an end-to-end framework that detects dysfluency with accurate type and time transcription, inspired by the YOLO object detection algorithm. Stutter-Solver can handle co-dysfluencies and is a natural multi-lingual dysfluency detector. To leverage scalability and boost performance, we also introduce three novel dysfluency corpora: VCTK-Pro, VCTK-Art, and AISHELL3-Pro, simulating natural spoken dysfluencies including repetition, block, missing, replacement, and prolongation through articulatory-encodec and TTS-based methods. Our approach achieves state-of-the-art performance on all available dysfluency corpora. Code and datasets are open-sourced at https://github.com/eureka235/Stutter-Solver |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_09621 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection Zhou, Xuanru Cho, Cheol Jun Sharma, Ayati Morin, Brittany Baquirin, David Vonk, Jet Ezzes, Zoe Miller, Zachary Tee, Boon Lead Tempini, Maria Luisa Gorno Lian, Jiachen Anumanchipalli, Gopala Audio and Speech Processing Artificial Intelligence Sound Current de-facto dysfluency modeling methods utilize template matching algorithms which are not generalizable to out-of-domain real-world dysfluencies across languages, and are not scalable with increasing amounts of training data. To handle these problems, we propose Stutter-Solver: an end-to-end framework that detects dysfluency with accurate type and time transcription, inspired by the YOLO object detection algorithm. Stutter-Solver can handle co-dysfluencies and is a natural multi-lingual dysfluency detector. To leverage scalability and boost performance, we also introduce three novel dysfluency corpora: VCTK-Pro, VCTK-Art, and AISHELL3-Pro, simulating natural spoken dysfluencies including repetition, block, missing, replacement, and prolongation through articulatory-encodec and TTS-based methods. Our approach achieves state-of-the-art performance on all available dysfluency corpora. Code and datasets are open-sourced at https://github.com/eureka235/Stutter-Solver |
| title | Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection |
| topic | Audio and Speech Processing Artificial Intelligence Sound |
| url | https://arxiv.org/abs/2409.09621 |