Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hono, Yukiya, Mitsuda, Koh, Zhao, Tianyu, Mitsui, Kentaro, Wakatsuki, Toshiaki, Sawada, Kei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911907095511040
author Hono, Yukiya
Mitsuda, Koh
Zhao, Tianyu
Mitsui, Kentaro
Wakatsuki, Toshiaki
Sawada, Kei
author_facet Hono, Yukiya
Mitsuda, Koh
Zhao, Tianyu
Mitsui, Kentaro
Wakatsuki, Toshiaki
Sawada, Kei
contents Advances in machine learning have made it possible to perform various text and speech processing tasks, such as automatic speech recognition (ASR), in an end-to-end (E2E) manner. E2E approaches utilizing pre-trained models are gaining attention for conserving training data and resources. However, most of their applications in ASR involve only one of either a pre-trained speech or a language model. This paper proposes integrating a pre-trained speech representation model and a large language model (LLM) for E2E ASR. The proposed model enables the optimization of the entire ASR process, including acoustic feature extraction and acoustic and language modeling, by combining pre-trained models with a bridge network and also enables the application of remarkable developments in LLM utilization, such as parameter-efficient domain adaptation and inference optimization. Experimental results demonstrate that the proposed model achieves a performance comparable to that of modern E2E ASR models by utilizing powerful pre-training models with the proposed integrated approach.
format Preprint
id arxiv_https___arxiv_org_abs_2312_03668
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
Hono, Yukiya
Mitsuda, Koh
Zhao, Tianyu
Mitsui, Kentaro
Wakatsuki, Toshiaki
Sawada, Kei
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Advances in machine learning have made it possible to perform various text and speech processing tasks, such as automatic speech recognition (ASR), in an end-to-end (E2E) manner. E2E approaches utilizing pre-trained models are gaining attention for conserving training data and resources. However, most of their applications in ASR involve only one of either a pre-trained speech or a language model. This paper proposes integrating a pre-trained speech representation model and a large language model (LLM) for E2E ASR. The proposed model enables the optimization of the entire ASR process, including acoustic feature extraction and acoustic and language modeling, by combining pre-trained models with a bridge network and also enables the application of remarkable developments in LLM utilization, such as parameter-efficient domain adaptation and inference optimization. Experimental results demonstrate that the proposed model achieves a performance comparable to that of modern E2E ASR models by utilizing powerful pre-training models with the proposed integrated approach.
title Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2312.03668