Multi-Scale Temporal Transformer For Speech Emotion Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhipeng, Xing, Xiaofen, Fang, Yuanbo, Zhang, Weibin, Fan, Hengsheng, Xu, Xiangmin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913525071347712
author Li, Zhipeng
Xing, Xiaofen
Fang, Yuanbo
Zhang, Weibin
Fan, Hengsheng
Xu, Xiangmin
author_facet Li, Zhipeng
Xing, Xiaofen
Fang, Yuanbo
Zhang, Weibin
Fan, Hengsheng
Xu, Xiangmin
contents Speech emotion recognition plays a crucial role in human-machine interaction systems. Recently various optimized Transformers have been successfully applied to speech emotion recognition. However, the existing Transformer architectures focus more on global information and require large computation. On the other hand, abundant speech emotional representations exist locally on different parts of the input speech. To tackle these problems, we propose a Multi-Scale TRansfomer (MSTR) for speech emotion recognition. It comprises of three main components: (1) a multi-scale temporal feature operator, (2) a fractal self-attention module, and (3) a scale mixer module. These three components can effectively enhance the transformer's ability to learn multi-scale local emotion representations. Experimental results demonstrate that the proposed MSTR model significantly outperforms a vanilla Transformer and other state-of-the-art methods across three speech emotion datasets: IEMOCAP, MELD and, CREMAD. In addition, it can greatly reduce the computational cost.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00390
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Scale Temporal Transformer For Speech Emotion Recognition
Li, Zhipeng
Xing, Xiaofen
Fang, Yuanbo
Zhang, Weibin
Fan, Hengsheng
Xu, Xiangmin
Audio and Speech Processing
Speech emotion recognition plays a crucial role in human-machine interaction systems. Recently various optimized Transformers have been successfully applied to speech emotion recognition. However, the existing Transformer architectures focus more on global information and require large computation. On the other hand, abundant speech emotional representations exist locally on different parts of the input speech. To tackle these problems, we propose a Multi-Scale TRansfomer (MSTR) for speech emotion recognition. It comprises of three main components: (1) a multi-scale temporal feature operator, (2) a fractal self-attention module, and (3) a scale mixer module. These three components can effectively enhance the transformer's ability to learn multi-scale local emotion representations. Experimental results demonstrate that the proposed MSTR model significantly outperforms a vanilla Transformer and other state-of-the-art methods across three speech emotion datasets: IEMOCAP, MELD and, CREMAD. In addition, it can greatly reduce the computational cost.
title Multi-Scale Temporal Transformer For Speech Emotion Recognition
topic Audio and Speech Processing
url https://arxiv.org/abs/2410.00390