SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lian, Jiachen, Zhou, Xuanru, Ezzes, Zoe, Vonk, Jet, Morin, Brittany, Baquirin, David, Mille, Zachary, Tempini, Maria Luisa Gorno, Anumanchipalli, Gopala Krishna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910722162688000
author Lian, Jiachen
Zhou, Xuanru
Ezzes, Zoe
Vonk, Jet
Morin, Brittany
Baquirin, David
Mille, Zachary
Tempini, Maria Luisa Gorno
Anumanchipalli, Gopala Krishna
author_facet Lian, Jiachen
Zhou, Xuanru
Ezzes, Zoe
Vonk, Jet
Morin, Brittany
Baquirin, David
Mille, Zachary
Tempini, Maria Luisa Gorno
Anumanchipalli, Gopala Krishna
contents Speech is a hierarchical collection of text, prosody, emotions, dysfluencies, etc. Automatic transcription of speech that goes beyond text (words) is an underexplored problem. We focus on transcribing speech along with non-fluencies (dysfluencies). The current state-of-the-art pipeline SSDM suffers from complex architecture design, training complexity, and significant shortcomings in the local sequence aligner, and it does not explore in-context learning capacity. In this work, we propose SSDM 2.0, which tackles those shortcomings via four main contributions: (1) We propose a novel \textit{neural articulatory flow} to derive highly scalable speech representations. (2) We developed a \textit{full-stack connectionist subsequence aligner} that captures all types of dysfluencies. (3) We introduced a mispronunciation prompt pipeline and consistency learning module into LLM to leverage dysfluency \textit{in-context pronunciation learning} abilities. (4) We curated Libri-Dys and open-sourced the current largest-scale co-dysfluency corpus, \textit{Libri-Co-Dys}, for future research endeavors. In clinical experiments on pathological speech transcription, we tested SSDM 2.0 using nfvPPA corpus primarily characterized by \textit{articulatory dysfluencies}. Overall, SSDM 2.0 outperforms SSDM and all other dysfluency transcription models by a large margin. See our project demo page at \url{https://berkeley-speech-group.github.io/SSDM2.0/}.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00265
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies
Lian, Jiachen
Zhou, Xuanru
Ezzes, Zoe
Vonk, Jet
Morin, Brittany
Baquirin, David
Mille, Zachary
Tempini, Maria Luisa Gorno
Anumanchipalli, Gopala Krishna
Audio and Speech Processing
Speech is a hierarchical collection of text, prosody, emotions, dysfluencies, etc. Automatic transcription of speech that goes beyond text (words) is an underexplored problem. We focus on transcribing speech along with non-fluencies (dysfluencies). The current state-of-the-art pipeline SSDM suffers from complex architecture design, training complexity, and significant shortcomings in the local sequence aligner, and it does not explore in-context learning capacity. In this work, we propose SSDM 2.0, which tackles those shortcomings via four main contributions: (1) We propose a novel \textit{neural articulatory flow} to derive highly scalable speech representations. (2) We developed a \textit{full-stack connectionist subsequence aligner} that captures all types of dysfluencies. (3) We introduced a mispronunciation prompt pipeline and consistency learning module into LLM to leverage dysfluency \textit{in-context pronunciation learning} abilities. (4) We curated Libri-Dys and open-sourced the current largest-scale co-dysfluency corpus, \textit{Libri-Co-Dys}, for future research endeavors. In clinical experiments on pathological speech transcription, we tested SSDM 2.0 using nfvPPA corpus primarily characterized by \textit{articulatory dysfluencies}. Overall, SSDM 2.0 outperforms SSDM and all other dysfluency transcription models by a large margin. See our project demo page at \url{https://berkeley-speech-group.github.io/SSDM2.0/}.
title SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies
topic Audio and Speech Processing
url https://arxiv.org/abs/2412.00265