Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yuhang, Yang, Yizhou, Peng, Chng, Eng Siong, Zhong, Xionghu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914955963400192
author Yuhang, Yang
Yizhou, Peng
Chng, Eng Siong
Zhong, Xionghu
author_facet Yuhang, Yang
Yizhou, Peng
Chng, Eng Siong
Zhong, Xionghu
contents The integration of large language models (LLMs) with pre-trained speech models has opened up new avenues in automatic speech recognition (ASR). While LLMs excel in multimodal understanding tasks, effectively leveraging their capabilities for ASR remains a significant challenge. This paper presents a novel training approach to enhance LLM performance in ASR tasks. We propose pre-training LLMs on Pinyin embedding sequences, which represent pronunciation features, to generate corresponding Chinese characters. This step enables the LLM to adapt to generating text from pronunciation features before encountering real speech data. Furthermore, we fine-tune the LoRA parameters to enhance the LLM's understanding of speech modality information. In AISHELL-1 corpus, our approach yields a 9.5% relative improvement in ASR tasks compared to the baseline without Pinyi-to-Character pre-training. Additionally, incorporating auxiliary text data for Pinyi-to-Character pre-training further boosts performance, achieving a 19.0% relative improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2409_16005
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
Yuhang, Yang
Yizhou, Peng
Chng, Eng Siong
Zhong, Xionghu
Computation and Language
Sound
Audio and Speech Processing
The integration of large language models (LLMs) with pre-trained speech models has opened up new avenues in automatic speech recognition (ASR). While LLMs excel in multimodal understanding tasks, effectively leveraging their capabilities for ASR remains a significant challenge. This paper presents a novel training approach to enhance LLM performance in ASR tasks. We propose pre-training LLMs on Pinyin embedding sequences, which represent pronunciation features, to generate corresponding Chinese characters. This step enables the LLM to adapt to generating text from pronunciation features before encountering real speech data. Furthermore, we fine-tune the LoRA parameters to enhance the LLM's understanding of speech modality information. In AISHELL-1 corpus, our approach yields a 9.5% relative improvement in ASR tasks compared to the baseline without Pinyi-to-Character pre-training. Additionally, incorporating auxiliary text data for Pinyi-to-Character pre-training further boosts performance, achieving a 19.0% relative improvement.
title Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.16005