Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Xuanru, Lian, Jiachen, Cho, Cheol Jun, Liu, Jingwen, Ye, Zongli, Zhang, Jinming, Morin, Brittany, Baquirin, David, Vonk, Jet, Ezzes, Zoe, Miller, Zachary, Tempini, Maria Luisa Gorno, Anumanchipalli, Gopala
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914953828499456
author Zhou, Xuanru
Lian, Jiachen
Cho, Cheol Jun
Liu, Jingwen
Ye, Zongli
Zhang, Jinming
Morin, Brittany
Baquirin, David
Vonk, Jet
Ezzes, Zoe
Miller, Zachary
Tempini, Maria Luisa Gorno
Anumanchipalli, Gopala
author_facet Zhou, Xuanru
Lian, Jiachen
Cho, Cheol Jun
Liu, Jingwen
Ye, Zongli
Zhang, Jinming
Morin, Brittany
Baquirin, David
Vonk, Jet
Ezzes, Zoe
Miller, Zachary
Tempini, Maria Luisa Gorno
Anumanchipalli, Gopala
contents Speech dysfluency modeling is a task to detect dysfluencies in speech, such as repetition, block, insertion, replacement, and deletion. Most recent advancements treat this problem as a time-based object detection problem. In this work, we revisit this problem from a new perspective: tokenizing dysfluencies and modeling the detection problem as a token-based automatic speech recognition (ASR) problem. We propose rule-based speech and text dysfluency simulators and develop VCTK-token, and then develop a Whisper-like seq2seq architecture to build a new benchmark with decent performance. We also systematically compare our proposed token-based methods with time-based methods, and propose a unified benchmark to facilitate future research endeavors. We open-source these resources for the broader scientific community. The project page is available at https://rorizzz.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2409_13582
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
Zhou, Xuanru
Lian, Jiachen
Cho, Cheol Jun
Liu, Jingwen
Ye, Zongli
Zhang, Jinming
Morin, Brittany
Baquirin, David
Vonk, Jet
Ezzes, Zoe
Miller, Zachary
Tempini, Maria Luisa Gorno
Anumanchipalli, Gopala
Audio and Speech Processing
Artificial Intelligence
Sound
Speech dysfluency modeling is a task to detect dysfluencies in speech, such as repetition, block, insertion, replacement, and deletion. Most recent advancements treat this problem as a time-based object detection problem. In this work, we revisit this problem from a new perspective: tokenizing dysfluencies and modeling the detection problem as a token-based automatic speech recognition (ASR) problem. We propose rule-based speech and text dysfluency simulators and develop VCTK-token, and then develop a Whisper-like seq2seq architecture to build a new benchmark with decent performance. We also systematically compare our proposed token-based methods with time-based methods, and propose a unified benchmark to facilitate future research endeavors. We open-source these resources for the broader scientific community. The project page is available at https://rorizzz.github.io/
title Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2409.13582