Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Geng, Xuelong, Xu, Tianyi, Wei, Kun, Mu, Bingshen, Xue, Hongfei, Wang, He, Li, Yangze, Guo, Pengcheng, Dai, Yuhang, Li, Longhao, Shao, Mingchen, Xie, Lei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909377426882560
author Geng, Xuelong
Xu, Tianyi
Wei, Kun
Mu, Bingshen
Xue, Hongfei
Wang, He
Li, Yangze
Guo, Pengcheng
Dai, Yuhang
Li, Longhao
Shao, Mingchen
Xie, Lei
author_facet Geng, Xuelong
Xu, Tianyi
Wei, Kun
Mu, Bingshen
Xue, Hongfei
Wang, He
Li, Yangze
Guo, Pengcheng
Dai, Yuhang
Li, Longhao
Shao, Mingchen
Xie, Lei
contents Large Language Models (LLMs) have demonstrated unparalleled effectiveness in various NLP tasks, and integrating LLMs with automatic speech recognition (ASR) is becoming a mainstream paradigm. Building upon this momentum, our research delves into an in-depth examination of this paradigm on a large open-source Chinese dataset. Specifically, our research aims to evaluate the impact of various configurations of speech encoders, LLMs, and projector modules in the context of the speech foundation encoder-LLM ASR paradigm. Furthermore, we introduce a three-stage training approach, expressly developed to enhance the model's ability to align auditory and textual information. The implementation of this approach, alongside the strategic integration of ASR components, enabled us to achieve the SOTA performance on the AISHELL-1, Test_Net, and Test_Meeting test sets. Our analysis presents an empirical foundation for future research in LLM-based ASR systems and offers insights into optimizing performance using Chinese datasets. We will publicly release all scripts used for data preparation, training, inference, and scoring, as well as pre-trained models and training logs to promote reproducible research.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02132
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
Geng, Xuelong
Xu, Tianyi
Wei, Kun
Mu, Bingshen
Xue, Hongfei
Wang, He
Li, Yangze
Guo, Pengcheng
Dai, Yuhang
Li, Longhao
Shao, Mingchen
Xie, Lei
Sound
Computation and Language
Audio and Speech Processing
Large Language Models (LLMs) have demonstrated unparalleled effectiveness in various NLP tasks, and integrating LLMs with automatic speech recognition (ASR) is becoming a mainstream paradigm. Building upon this momentum, our research delves into an in-depth examination of this paradigm on a large open-source Chinese dataset. Specifically, our research aims to evaluate the impact of various configurations of speech encoders, LLMs, and projector modules in the context of the speech foundation encoder-LLM ASR paradigm. Furthermore, we introduce a three-stage training approach, expressly developed to enhance the model's ability to align auditory and textual information. The implementation of this approach, alongside the strategic integration of ASR components, enabled us to achieve the SOTA performance on the AISHELL-1, Test_Net, and Test_Meeting test sets. Our analysis presents an empirical foundation for future research in LLM-based ASR systems and offers insights into optimizing performance using Chinese datasets. We will publicly release all scripts used for data preparation, training, inference, and scoring, as well as pre-trained models and training logs to promote reproducible research.
title Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2405.02132