R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910393597689856 |
|---|---|
| author | Chang, Heng-Jui Glass, James |
| author_facet | Chang, Heng-Jui Glass, James |
| contents | This paper introduces Robust Spin (R-Spin), a data-efficient domain-specific self-supervision method for speaker and noise-invariant speech representations by learning discrete acoustic units with speaker-invariant clustering (Spin). R-Spin resolves Spin's issues and enhances content representations by learning to predict acoustic pieces. R-Spin offers a 12X reduction in computational resources compared to previous state-of-the-art methods while outperforming them in severely distorted speech scenarios. This paper provides detailed analyses to show how discrete units contribute to speech encoder training and improving robustness in diverse acoustic environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_09117 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces Chang, Heng-Jui Glass, James Computation and Language Sound Audio and Speech Processing This paper introduces Robust Spin (R-Spin), a data-efficient domain-specific self-supervision method for speaker and noise-invariant speech representations by learning discrete acoustic units with speaker-invariant clustering (Spin). R-Spin resolves Spin's issues and enhances content representations by learning to predict acoustic pieces. R-Spin offers a 12X reduction in computational resources compared to previous state-of-the-art methods while outperforming them in severely distorted speech scenarios. This paper provides detailed analyses to show how discrete units contribute to speech encoder training and improving robustness in diverse acoustic environments. |
| title | R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2311.09117 |