NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, He, Park, Taejin, Dhawan, Kunal, Medennikov, Ivan, Puvvada, Krishna C., Koluguri, Nithin Rao, Wang, Weiqing, Balam, Jagadeesh, Ginsburg, Boris
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913655225843712
author Huang, He
Park, Taejin
Dhawan, Kunal
Medennikov, Ivan
Puvvada, Krishna C.
Koluguri, Nithin Rao
Wang, Weiqing
Balam, Jagadeesh
Ginsburg, Boris
author_facet Huang, He
Park, Taejin
Dhawan, Kunal
Medennikov, Ivan
Puvvada, Krishna C.
Koluguri, Nithin Rao
Wang, Weiqing
Balam, Jagadeesh
Ginsburg, Boris
contents Self-supervised learning has been proved to benefit a wide range of speech processing tasks, such as speech recognition/translation, speaker verification and diarization, etc. However, most of current approaches are computationally expensive. In this paper, we propose a simplified and more efficient self-supervised learning framework termed as NeMo Encoder for Speech Tasks (NEST). Specifically, we adopt the FastConformer architecture with 8x sub-sampling rate, which is faster than Transformer or Conformer architectures. Instead of clustering-based quantization, we use fixed random projection for its simplicity and effectiveness. We also implement a generalized noisy speech augmentation that teaches the model to disentangle the main speaker from noise or other speakers. Experiments show that \model improves over existing self-supervised models and achieves new state-of-the-art performance on a variety of speech processing tasks, such as speech recognition/translation, speaker diarization, spoken language understanding, etc. Code and checkpoints are publicly available via NVIDIA NeMo framework.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13106
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
Huang, He
Park, Taejin
Dhawan, Kunal
Medennikov, Ivan
Puvvada, Krishna C.
Koluguri, Nithin Rao
Wang, Weiqing
Balam, Jagadeesh
Ginsburg, Boris
Sound
Audio and Speech Processing
Self-supervised learning has been proved to benefit a wide range of speech processing tasks, such as speech recognition/translation, speaker verification and diarization, etc. However, most of current approaches are computationally expensive. In this paper, we propose a simplified and more efficient self-supervised learning framework termed as NeMo Encoder for Speech Tasks (NEST). Specifically, we adopt the FastConformer architecture with 8x sub-sampling rate, which is faster than Transformer or Conformer architectures. Instead of clustering-based quantization, we use fixed random projection for its simplicity and effectiveness. We also implement a generalized noisy speech augmentation that teaches the model to disentangle the main speaker from noise or other speakers. Experiments show that \model improves over existing self-supervised models and achieves new state-of-the-art performance on a variety of speech processing tasks, such as speech recognition/translation, speaker diarization, spoken language understanding, etc. Code and checkpoints are publicly available via NVIDIA NeMo framework.
title NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2408.13106