FGCL: Fine-grained Contrastive Learning For Mandarin Stuttering Event Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Han, Wang, Wenyu, Zhou, Yiquan, Ding, Hongwu, Xu, Jiacheng, Zhu, Jihua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913536016384000
author Jiang, Han
Wang, Wenyu
Zhou, Yiquan
Ding, Hongwu
Xu, Jiacheng
Zhu, Jihua
author_facet Jiang, Han
Wang, Wenyu
Zhou, Yiquan
Ding, Hongwu
Xu, Jiacheng
Zhu, Jihua
contents This paper presents the T031 team's approach to the StutteringSpeech Challenge in SLT2024. Mandarin Stuttering Event Detection (MSED) aims to detect instances of stuttering events in Mandarin speech. We propose a detailed acoustic analysis method to improve the accuracy of stutter detection by capturing subtle nuances that previous Stuttering Event Detection (SED) techniques have overlooked. To this end, we introduce the Fine-Grained Contrastive Learning (FGCL) framework for MSED. Specifically, we model the frame-level probabilities of stuttering events and introduce a mining algorithm to identify both easy and confusing frames. Then, we propose a stutter contrast loss to enhance the distinction between stuttered and fluent speech frames, thereby improving the discriminative capability of stuttered feature embeddings. Extensive evaluations on English and Mandarin datasets demonstrate the effectiveness of FGCL, achieving a significant increase of over 5.0% in F1 score on Mandarin data.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05647
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FGCL: Fine-grained Contrastive Learning For Mandarin Stuttering Event Detection
Jiang, Han
Wang, Wenyu
Zhou, Yiquan
Ding, Hongwu
Xu, Jiacheng
Zhu, Jihua
Sound
Audio and Speech Processing
This paper presents the T031 team's approach to the StutteringSpeech Challenge in SLT2024. Mandarin Stuttering Event Detection (MSED) aims to detect instances of stuttering events in Mandarin speech. We propose a detailed acoustic analysis method to improve the accuracy of stutter detection by capturing subtle nuances that previous Stuttering Event Detection (SED) techniques have overlooked. To this end, we introduce the Fine-Grained Contrastive Learning (FGCL) framework for MSED. Specifically, we model the frame-level probabilities of stuttering events and introduce a mining algorithm to identify both easy and confusing frames. Then, we propose a stutter contrast loss to enhance the distinction between stuttered and fluent speech frames, thereby improving the discriminative capability of stuttered feature embeddings. Extensive evaluations on English and Mandarin datasets demonstrate the effectiveness of FGCL, achieving a significant increase of over 5.0% in F1 score on Mandarin data.
title FGCL: Fine-grained Contrastive Learning For Mandarin Stuttering Event Detection
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.05647