Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Ziyang, Wu, Wen, Zheng, Zhisheng, Guo, Yiwei, Chen, Qian, Zhang, Shiliang, Chen, Xie
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916455384088576
author Ma, Ziyang
Wu, Wen
Zheng, Zhisheng
Guo, Yiwei
Chen, Qian
Zhang, Shiliang
Chen, Xie
author_facet Ma, Ziyang
Wu, Wen
Zheng, Zhisheng
Guo, Yiwei
Chen, Qian
Zhang, Shiliang
Chen, Xie
contents In this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. First, we investigated the representation ability of different speech self-supervised pre-trained models, and we found that data2vec has a good representation ability on the SER task. Second, we employed a powerful large language model (LLM), GPT-4, and emotional text-to-speech (TTS) model, Azure TTS, to generate emotionally congruent text and speech. We carefully designed the text prompt and dataset construction, to obtain the synthetic emotional speech data with high quality. Third, we studied different ways of data augmentation to promote the SER task with synthetic speech, including random mixing, adversarial training, transfer learning, and curriculum learning. Experiments and ablation studies on the IEMOCAP dataset demonstrate the effectiveness of our method, compared with other data augmentation methods, and data augmentation with other synthetic data.
format Preprint
id arxiv_https___arxiv_org_abs_2309_10294
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
Ma, Ziyang
Wu, Wen
Zheng, Zhisheng
Guo, Yiwei
Chen, Qian
Zhang, Shiliang
Chen, Xie
Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
In this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. First, we investigated the representation ability of different speech self-supervised pre-trained models, and we found that data2vec has a good representation ability on the SER task. Second, we employed a powerful large language model (LLM), GPT-4, and emotional text-to-speech (TTS) model, Azure TTS, to generate emotionally congruent text and speech. We carefully designed the text prompt and dataset construction, to obtain the synthetic emotional speech data with high quality. Third, we studied different ways of data augmentation to promote the SER task with synthetic speech, including random mixing, adversarial training, transfer learning, and curriculum learning. Experiments and ablation studies on the IEMOCAP dataset demonstrate the effectiveness of our method, compared with other data augmentation methods, and data augmentation with other synthetic data.
title Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
topic Computation and Language
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2309.10294