State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Farhadipour, Aref, Beigi, Homayoon, Dellwo, Volker, Veisi, Hadi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911025616388096
author Farhadipour, Aref
Beigi, Homayoon
Dellwo, Volker
Veisi, Hadi
author_facet Farhadipour, Aref
Beigi, Homayoon
Dellwo, Volker
Veisi, Hadi
contents Whispered speech recognition presents significant challenges for conventional automatic speech recognition systems, particularly when combined with dialect variation. However, utilizing an efficient method to solve this problem using a low-range dataset and processing load is beneficial. This paper proposes a solution using a Mamba-based state-space model and four fine-tuned self-supervised models consisting of Wav2Vec2, WavLM, HuBERT, and Whisper to address the dual challenges of whispered speech and dialect diversity. Based on our knowledge, this represents the best performance reported on the wTIMIT and CHAINS datasets for whispered speech recognition. We trained the models using whispered and normal speech data across Singaporean, US, and Irish dialects. The findings demonstrated that utilizing the proposed Mamba-based model could work as a highly efficient model trained with low amounts of whispered data to simultaneously work on whispered and normal speech recognition. The code for this work is freely available.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16969
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
Farhadipour, Aref
Beigi, Homayoon
Dellwo, Volker
Veisi, Hadi
Audio and Speech Processing
Sound
Whispered speech recognition presents significant challenges for conventional automatic speech recognition systems, particularly when combined with dialect variation. However, utilizing an efficient method to solve this problem using a low-range dataset and processing load is beneficial. This paper proposes a solution using a Mamba-based state-space model and four fine-tuned self-supervised models consisting of Wav2Vec2, WavLM, HuBERT, and Whisper to address the dual challenges of whispered speech and dialect diversity. Based on our knowledge, this represents the best performance reported on the wTIMIT and CHAINS datasets for whispered speech recognition. We trained the models using whispered and normal speech data across Singaporean, US, and Irish dialects. The findings demonstrated that utilizing the proposed Mamba-based model could work as a highly efficient model trained with low amounts of whispered data to simultaneously work on whispered and normal speech recognition. The code for this work is freely available.
title State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.16969