BLSP-Emo: Towards Empathetic Large Speech-Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Chen, Liao, Minpeng, Huang, Zhongqiang, Wu, Junhong, Zong, Chengqing, Zhang, Jiajun
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911908070686720
author Wang, Chen
Liao, Minpeng
Huang, Zhongqiang
Wu, Junhong
Zong, Chengqing
Zhang, Jiajun
author_facet Wang, Chen
Liao, Minpeng
Huang, Zhongqiang
Wu, Junhong
Zong, Chengqing
Zhang, Jiajun
contents The recent release of GPT-4o showcased the potential of end-to-end multimodal models, not just in terms of low latency but also in their ability to understand and generate expressive speech with rich emotions. While the details are unknown to the open research community, it likely involves significant amounts of curated data and compute, neither of which is readily accessible. In this paper, we present BLSP-Emo (Bootstrapped Language-Speech Pretraining with Emotion support), a novel approach to developing an end-to-end speech-language model capable of understanding both semantics and emotions in speech and generate empathetic responses. BLSP-Emo utilizes existing speech recognition (ASR) and speech emotion recognition (SER) datasets through a two-stage process. The first stage focuses on semantic alignment, following recent work on pretraining speech-language models using ASR data. The second stage performs emotion alignment with the pretrained speech-language model on an emotion-aware continuation task constructed from SER data. Our experiments demonstrate that the BLSP-Emo model excels in comprehending speech and delivering empathetic responses, both in instruction-following tasks and conversations.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03872
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BLSP-Emo: Towards Empathetic Large Speech-Language Models
Wang, Chen
Liao, Minpeng
Huang, Zhongqiang
Wu, Junhong
Zong, Chengqing
Zhang, Jiajun
Computation and Language
Sound
Audio and Speech Processing
The recent release of GPT-4o showcased the potential of end-to-end multimodal models, not just in terms of low latency but also in their ability to understand and generate expressive speech with rich emotions. While the details are unknown to the open research community, it likely involves significant amounts of curated data and compute, neither of which is readily accessible. In this paper, we present BLSP-Emo (Bootstrapped Language-Speech Pretraining with Emotion support), a novel approach to developing an end-to-end speech-language model capable of understanding both semantics and emotions in speech and generate empathetic responses. BLSP-Emo utilizes existing speech recognition (ASR) and speech emotion recognition (SER) datasets through a two-stage process. The first stage focuses on semantic alignment, following recent work on pretraining speech-language models using ASR data. The second stage performs emotion alignment with the pretrained speech-language model on an emotion-aware continuation task constructed from SER data. Our experiments demonstrate that the BLSP-Emo model excels in comprehending speech and delivering empathetic responses, both in instruction-following tasks and conversations.
title BLSP-Emo: Towards Empathetic Large Speech-Language Models
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.03872