Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Weidong, Xing, Xiaofen, Chen, Peihao, Xu, Xiangmin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914758710525952
author Chen, Weidong
Xing, Xiaofen
Chen, Peihao
Xu, Xiangmin
author_facet Chen, Weidong
Xing, Xiaofen
Chen, Peihao
Xu, Xiangmin
contents This paper presents a paradigm that adapts general large-scale pretrained models (PTMs) to speech emotion recognition task. Although PTMs shed new light on artificial general intelligence, they are constructed with general tasks in mind, and thus, their efficacy for specific tasks can be further improved. Additionally, employing PTMs in practical applications can be challenging due to their considerable size. Above limitations spawn another research direction, namely, optimizing large-scale PTMs for specific tasks to generate task-specific PTMs that are both compact and effective. In this paper, we focus on the speech emotion recognition task and propose an improved emotion-specific pretrained encoder called Vesper. Vesper is pretrained on a speech dataset based on WavLM and takes into account emotional characteristics. To enhance sensitivity to emotional information, Vesper employs an emotion-guided masking strategy to identify the regions that need masking. Subsequently, Vesper employs hierarchical and cross-layer self-supervision to improve its ability to capture acoustic and semantic representations, both of which are crucial for emotion recognition. Experimental results on the IEMOCAP, MELD, and CREMA-D datasets demonstrate that Vesper with 4 layers outperforms WavLM Base with 12 layers, and the performance of Vesper with 12 layers surpasses that of WavLM Large with 24 layers.
format Preprint
id arxiv_https___arxiv_org_abs_2307_10757
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition
Chen, Weidong
Xing, Xiaofen
Chen, Peihao
Xu, Xiangmin
Sound
Computation and Language
Audio and Speech Processing
This paper presents a paradigm that adapts general large-scale pretrained models (PTMs) to speech emotion recognition task. Although PTMs shed new light on artificial general intelligence, they are constructed with general tasks in mind, and thus, their efficacy for specific tasks can be further improved. Additionally, employing PTMs in practical applications can be challenging due to their considerable size. Above limitations spawn another research direction, namely, optimizing large-scale PTMs for specific tasks to generate task-specific PTMs that are both compact and effective. In this paper, we focus on the speech emotion recognition task and propose an improved emotion-specific pretrained encoder called Vesper. Vesper is pretrained on a speech dataset based on WavLM and takes into account emotional characteristics. To enhance sensitivity to emotional information, Vesper employs an emotion-guided masking strategy to identify the regions that need masking. Subsequently, Vesper employs hierarchical and cross-layer self-supervision to improve its ability to capture acoustic and semantic representations, both of which are crucial for emotion recognition. Experimental results on the IEMOCAP, MELD, and CREMA-D datasets demonstrate that Vesper with 4 layers outperforms WavLM Base with 12 layers, and the performance of Vesper with 12 layers surpasses that of WavLM Large with 24 layers.
title Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2307.10757