R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chang, Heng-Jui, Glass, James
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910393597689856
author Chang, Heng-Jui
Glass, James
author_facet Chang, Heng-Jui
Glass, James
contents This paper introduces Robust Spin (R-Spin), a data-efficient domain-specific self-supervision method for speaker and noise-invariant speech representations by learning discrete acoustic units with speaker-invariant clustering (Spin). R-Spin resolves Spin's issues and enhances content representations by learning to predict acoustic pieces. R-Spin offers a 12X reduction in computational resources compared to previous state-of-the-art methods while outperforming them in severely distorted speech scenarios. This paper provides detailed analyses to show how discrete units contribute to speech encoder training and improving robustness in diverse acoustic environments.
format Preprint
id arxiv_https___arxiv_org_abs_2311_09117
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces
Chang, Heng-Jui
Glass, James
Computation and Language
Sound
Audio and Speech Processing
This paper introduces Robust Spin (R-Spin), a data-efficient domain-specific self-supervision method for speaker and noise-invariant speech representations by learning discrete acoustic units with speaker-invariant clustering (Spin). R-Spin resolves Spin's issues and enhances content representations by learning to predict acoustic pieces. R-Spin offers a 12X reduction in computational resources compared to previous state-of-the-art methods while outperforming them in severely distorted speech scenarios. This paper provides detailed analyses to show how discrete units contribute to speech encoder training and improving robustness in diverse acoustic environments.
title R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2311.09117