Saved in:
Bibliographic Details
Main Authors: Li, Yuang, Yu, Jiawei, Zhang, Min, Ren, Mengxin, Zhao, Yanqing, Zhao, Xiaofeng, Tao, Shimin, Su, Jinsong, Yang, Hao
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2401.11382
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929375726796800
author Li, Yuang
Yu, Jiawei
Zhang, Min
Ren, Mengxin
Zhao, Yanqing
Zhao, Xiaofeng
Tao, Shimin
Su, Jinsong
Yang, Hao
author_facet Li, Yuang
Yu, Jiawei
Zhang, Min
Ren, Mengxin
Zhao, Yanqing
Zhao, Xiaofeng
Tao, Shimin
Su, Jinsong
Yang, Hao
contents Mapping speech tokens to the same feature space as text tokens has become the paradigm for the integration of speech modality into decoder-only large language models (LLMs). An alternative approach is to use an encoder-decoder architecture that incorporates speech features through cross-attention. This approach, however, has received less attention in the literature. In this work, we connect the Whisper encoder with ChatGLM3 and provide in-depth comparisons of these two approaches using Chinese automatic speech recognition (ASR) and name entity recognition (NER) tasks. We evaluate them not only by conventional metrics like the F1 score but also by a novel fine-grained taxonomy of ASR-NER errors. Our experiments reveal that encoder-decoder architecture outperforms decoder-only architecture with a short context, while decoder-only architecture benefits from a long context as it fully exploits all layers of the LLM. By using LLM, we significantly reduced the entity omission errors and improved the entity ASR accuracy compared to the Conformer baseline. Additionally, we obtained a state-of-the-art (SOTA) F1 score of 0.805 on the AISHELL-NER test set by using chain-of-thought (CoT) NER which first infers long-form ASR transcriptions and then predicts NER labels.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11382
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Using Large Language Model for End-to-End Chinese ASR and NER
Li, Yuang
Yu, Jiawei
Zhang, Min
Ren, Mengxin
Zhao, Yanqing
Zhao, Xiaofeng
Tao, Shimin
Su, Jinsong
Yang, Hao
Computation and Language
Artificial Intelligence
Mapping speech tokens to the same feature space as text tokens has become the paradigm for the integration of speech modality into decoder-only large language models (LLMs). An alternative approach is to use an encoder-decoder architecture that incorporates speech features through cross-attention. This approach, however, has received less attention in the literature. In this work, we connect the Whisper encoder with ChatGLM3 and provide in-depth comparisons of these two approaches using Chinese automatic speech recognition (ASR) and name entity recognition (NER) tasks. We evaluate them not only by conventional metrics like the F1 score but also by a novel fine-grained taxonomy of ASR-NER errors. Our experiments reveal that encoder-decoder architecture outperforms decoder-only architecture with a short context, while decoder-only architecture benefits from a long context as it fully exploits all layers of the LLM. By using LLM, we significantly reduced the entity omission errors and improved the entity ASR accuracy compared to the Conformer baseline. Additionally, we obtained a state-of-the-art (SOTA) F1 score of 0.805 on the AISHELL-NER test set by using chain-of-thought (CoT) NER which first infers long-form ASR transcriptions and then predicts NER labels.
title Using Large Language Model for End-to-End Chinese ASR and NER
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.11382