Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Meng, Lingwei, Hu, Shujie, Kang, Jiawen, Li, Zhaoqing, Wang, Yuejiao, Wu, Wenxuan, Wu, Xixin, Liu, Xunying, Meng, Helen
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909561682657280
author Meng, Lingwei
Hu, Shujie
Kang, Jiawen
Li, Zhaoqing
Wang, Yuejiao
Wu, Wenxuan
Wu, Xixin
Liu, Xunying
Meng, Helen
author_facet Meng, Lingwei
Hu, Shujie
Kang, Jiawen
Li, Zhaoqing
Wang, Yuejiao
Wu, Wenxuan
Wu, Xixin
Liu, Xunying
Meng, Helen
contents Recent advancements in large language models (LLMs) have revolutionized various domains, bringing significant progress and new opportunities. Despite progress in speech-related tasks, LLMs have not been sufficiently explored in multi-talker scenarios. In this work, we present a pioneering effort to investigate the capability of LLMs in transcribing speech in multi-talker environments, following versatile instructions related to multi-talker automatic speech recognition (ASR), target talker ASR, and ASR based on specific talker attributes such as sex, occurrence order, language, and keyword spoken. Our approach utilizes WavLM and Whisper encoder to extract multi-faceted speech representations that are sensitive to speaker characteristics and semantic context. These representations are then fed into an LLM fine-tuned using LoRA, enabling the capabilities for speech comprehension and transcription. Comprehensive experiments reveal the promising performance of our proposed system, MT-LLM, in cocktail party scenarios, highlighting the potential of LLM to handle speech-related tasks based on user instructions in such complex settings. The code, model, and samples are available at https://github.com/cuhealthybrains/MT-LLM.
format Preprint
id arxiv_https___arxiv_org_abs_2409_08596
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
Meng, Lingwei
Hu, Shujie
Kang, Jiawen
Li, Zhaoqing
Wang, Yuejiao
Wu, Wenxuan
Wu, Xixin
Liu, Xunying
Meng, Helen
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
Recent advancements in large language models (LLMs) have revolutionized various domains, bringing significant progress and new opportunities. Despite progress in speech-related tasks, LLMs have not been sufficiently explored in multi-talker scenarios. In this work, we present a pioneering effort to investigate the capability of LLMs in transcribing speech in multi-talker environments, following versatile instructions related to multi-talker automatic speech recognition (ASR), target talker ASR, and ASR based on specific talker attributes such as sex, occurrence order, language, and keyword spoken. Our approach utilizes WavLM and Whisper encoder to extract multi-faceted speech representations that are sensitive to speaker characteristics and semantic context. These representations are then fed into an LLM fine-tuned using LoRA, enabling the capabilities for speech comprehension and transcription. Comprehensive experiments reveal the promising performance of our proposed system, MT-LLM, in cocktail party scenarios, highlighting the potential of LLM to handle speech-related tasks based on user instructions in such complex settings. The code, model, and samples are available at https://github.com/cuhealthybrains/MT-LLM.
title Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.08596