Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Shih-heng, Shi, Jiatong, Huang, Chien-yu, Watanabe, Shinji, Lee, Hung-yi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916497124753408
author Wang, Shih-heng
Shi, Jiatong
Huang, Chien-yu
Watanabe, Shinji
Lee, Hung-yi
author_facet Wang, Shih-heng
Shi, Jiatong
Huang, Chien-yu
Watanabe, Shinji
Lee, Hung-yi
contents Self-supervised learning (SSL) models have shown exceptional capabilities across various speech-processing tasks. Continuous SSL representations are effective but suffer from high computational and storage demands. On the other hand, discrete SSL representations, although with degraded performance, reduce transmission and storage costs, and improve input sequence efficiency through de-duplication and subword-modeling. To boost the performance of discrete representations for ASR, we introduce a novel fusion mechanism that integrates two discrete representations. The fusion mechanism preserves all the benefits of discrete representation while enhancing the model's performance by integrating complementary information. Additionally, we explore "self-augmented'' discrete representations, which apply transformations to a single continuous SSL representation, eliminating the fusion mechanism's dependency on multiple SSL models and further decreasing its inference costs. Experimental results on benchmarks, including LibriSpeech and ML-SUPERB, indicate up to 19% and 24% relative character error rate improvement compared with the non-fusion baseline, validating the effectiveness of our proposed methods.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18107
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
Wang, Shih-heng
Shi, Jiatong
Huang, Chien-yu
Watanabe, Shinji
Lee, Hung-yi
Sound
Audio and Speech Processing
Self-supervised learning (SSL) models have shown exceptional capabilities across various speech-processing tasks. Continuous SSL representations are effective but suffer from high computational and storage demands. On the other hand, discrete SSL representations, although with degraded performance, reduce transmission and storage costs, and improve input sequence efficiency through de-duplication and subword-modeling. To boost the performance of discrete representations for ASR, we introduce a novel fusion mechanism that integrates two discrete representations. The fusion mechanism preserves all the benefits of discrete representation while enhancing the model's performance by integrating complementary information. Additionally, we explore "self-augmented'' discrete representations, which apply transformations to a single continuous SSL representation, eliminating the fusion mechanism's dependency on multiple SSL models and further decreasing its inference costs. Experimental results on benchmarks, including LibriSpeech and ML-SUPERB, indicate up to 19% and 24% relative character error rate improvement compared with the non-fusion baseline, validating the effectiveness of our proposed methods.
title Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2411.18107