DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lu, Ke-Han, Chen, Zhehuai, Fu, Szu-Wei, Yang, Chao-Han Huck, Balam, Jagadeesh, Ginsburg, Boris, Wang, Yu-Chiang Frank, Lee, Hung-yi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908469734408192
author Lu, Ke-Han
Chen, Zhehuai
Fu, Szu-Wei
Yang, Chao-Han Huck
Balam, Jagadeesh
Ginsburg, Boris
Wang, Yu-Chiang Frank
Lee, Hung-yi
author_facet Lu, Ke-Han
Chen, Zhehuai
Fu, Szu-Wei
Yang, Chao-Han Huck
Balam, Jagadeesh
Ginsburg, Boris
Wang, Yu-Chiang Frank
Lee, Hung-yi
contents Recent end-to-end speech language models (SLMs) have expanded upon the capabilities of large language models (LLMs) by incorporating pre-trained speech models. However, these SLMs often undergo extensive speech instruction-tuning to bridge the gap between speech and text modalities. This requires significant annotation efforts and risks catastrophic forgetting of the original language capabilities. In this work, we present a simple yet effective automatic process for creating speech-text pair data that carefully injects speech paralinguistic understanding abilities into SLMs while preserving the inherent language capabilities of the text-based LLM. Our model demonstrates general capabilities for speech-related tasks without the need for speech instruction-tuning data, achieving impressive performance on Dynamic-SUPERB and AIR-Bench-Chat benchmarks. Furthermore, our model exhibits the ability to follow complex instructions derived from LLMs, such as specific output formatting and chain-of-thought reasoning. Our approach not only enhances the versatility and effectiveness of SLMs but also reduces reliance on extensive annotated datasets, paving the way for more efficient and capable speech understanding systems.
format Preprint
id arxiv_https___arxiv_org_abs_2409_20007
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
Lu, Ke-Han
Chen, Zhehuai
Fu, Szu-Wei
Yang, Chao-Han Huck
Balam, Jagadeesh
Ginsburg, Boris
Wang, Yu-Chiang Frank
Lee, Hung-yi
Audio and Speech Processing
Computation and Language
Sound
Recent end-to-end speech language models (SLMs) have expanded upon the capabilities of large language models (LLMs) by incorporating pre-trained speech models. However, these SLMs often undergo extensive speech instruction-tuning to bridge the gap between speech and text modalities. This requires significant annotation efforts and risks catastrophic forgetting of the original language capabilities. In this work, we present a simple yet effective automatic process for creating speech-text pair data that carefully injects speech paralinguistic understanding abilities into SLMs while preserving the inherent language capabilities of the text-based LLM. Our model demonstrates general capabilities for speech-related tasks without the need for speech instruction-tuning data, achieving impressive performance on Dynamic-SUPERB and AIR-Bench-Chat benchmarks. Furthermore, our model exhibits the ability to follow complex instructions derived from LLMs, such as specific output formatting and chain-of-thought reasoning. Our approach not only enhances the versatility and effectiveness of SLMs but also reduces reliance on extensive annotated datasets, paving the way for more efficient and capable speech understanding systems.
title DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2409.20007