ASTRA: Aligning Speech and Text Representations for Asr without Sampling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gaur, Neeraj, Agrawal, Rohan, Wang, Gary, Haghani, Parisa, Rosenberg, Andrew, Ramabhadran, Bhuvana
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912046915780608
author Gaur, Neeraj
Agrawal, Rohan
Wang, Gary
Haghani, Parisa
Rosenberg, Andrew
Ramabhadran, Bhuvana
author_facet Gaur, Neeraj
Agrawal, Rohan
Wang, Gary
Haghani, Parisa
Rosenberg, Andrew
Ramabhadran, Bhuvana
contents This paper introduces ASTRA, a novel method for improving Automatic Speech Recognition (ASR) through text injection.Unlike prevailing techniques, ASTRA eliminates the need for sampling to match sequence lengths between speech and text modalities. Instead, it leverages the inherent alignments learned within CTC/RNNT models. This approach offers the following two advantages, namely, avoiding potential misalignment between speech and text features that could arise from upsampling and eliminating the need for models to accurately predict duration of sub-word tokens. This novel formulation of modality (length) matching as a weighted RNNT objective matches the performance of the state-of-the-art duration-based methods on the FLEURS benchmark, while opening up other avenues of research in speech processing.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06664
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ASTRA: Aligning Speech and Text Representations for Asr without Sampling
Gaur, Neeraj
Agrawal, Rohan
Wang, Gary
Haghani, Parisa
Rosenberg, Andrew
Ramabhadran, Bhuvana
Audio and Speech Processing
Machine Learning
Sound
This paper introduces ASTRA, a novel method for improving Automatic Speech Recognition (ASR) through text injection.Unlike prevailing techniques, ASTRA eliminates the need for sampling to match sequence lengths between speech and text modalities. Instead, it leverages the inherent alignments learned within CTC/RNNT models. This approach offers the following two advantages, namely, avoiding potential misalignment between speech and text features that could arise from upsampling and eliminating the need for models to accurately predict duration of sub-word tokens. This novel formulation of modality (length) matching as a weighted RNNT objective matches the performance of the state-of-the-art duration-based methods on the FLEURS benchmark, while opening up other avenues of research in speech processing.
title ASTRA: Aligning Speech and Text Representations for Asr without Sampling
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2406.06664