Comparing Discrete and Continuous Space LLMs for Speech Recognition

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Yaoxun, Zhang, Shi-Xiong, Yu, Jianwei, Wu, Zhiyong, Yu, Dong
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912010695868416
author Xu, Yaoxun
Zhang, Shi-Xiong
Yu, Jianwei
Wu, Zhiyong
Yu, Dong
author_facet Xu, Yaoxun
Zhang, Shi-Xiong
Yu, Jianwei
Wu, Zhiyong
Yu, Dong
contents This paper investigates discrete and continuous speech representations in Large Language Model (LLM)-based Automatic Speech Recognition (ASR), organizing them by feature continuity and training approach into four categories: supervised and unsupervised for both discrete and continuous types. We further classify LLMs based on their input and autoregressive feedback into continuous and discrete-space models. Using specialized encoders and comparative analysis with a Joint-Training-From-Scratch Language Model (JTFS LM) and pre-trained LLaMA2-7b, we provide a detailed examination of their effectiveness. Our work marks the first extensive comparison of speech representations in LLM-based ASR and explores various modeling techniques. We present an open-sourced achievement of a state-of-the-art Word Error Rate (WER) of 1.69\% on LibriSpeech using a HuBERT encoder, offering valuable insights for advancing ASR and natural language processing (NLP) research.
format Preprint
id arxiv_https___arxiv_org_abs_2409_00800
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Comparing Discrete and Continuous Space LLMs for Speech Recognition
Xu, Yaoxun
Zhang, Shi-Xiong
Yu, Jianwei
Wu, Zhiyong
Yu, Dong
Computation and Language
This paper investigates discrete and continuous speech representations in Large Language Model (LLM)-based Automatic Speech Recognition (ASR), organizing them by feature continuity and training approach into four categories: supervised and unsupervised for both discrete and continuous types. We further classify LLMs based on their input and autoregressive feedback into continuous and discrete-space models. Using specialized encoders and comparative analysis with a Joint-Training-From-Scratch Language Model (JTFS LM) and pre-trained LLaMA2-7b, we provide a detailed examination of their effectiveness. Our work marks the first extensive comparison of speech representations in LLM-based ASR and explores various modeling techniques. We present an open-sourced achievement of a state-of-the-art Word Error Rate (WER) of 1.69\% on LibriSpeech using a HuBERT encoder, offering valuable insights for advancing ASR and natural language processing (NLP) research.
title Comparing Discrete and Continuous Space LLMs for Speech Recognition
topic Computation and Language
url https://arxiv.org/abs/2409.00800