Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hyeon, Jonghwan, Oh, Yung-Hwan, Choi, Ho-Jin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910652707110912
author Hyeon, Jonghwan
Oh, Yung-Hwan
Choi, Ho-Jin
author_facet Hyeon, Jonghwan
Oh, Yung-Hwan
Choi, Ho-Jin
contents Speech Emotion Recognition (SER) analyzes human emotions expressed through speech. Self-supervised learning (SSL) offers a promising approach to SER by learning meaningful representations from a large amount of unlabeled audio data. However, existing SSL-based methods rely on Global Average Pooling (GAP) to represent audio signals, treating speech and non-speech segments equally. This can lead to dilution of informative speech features by irrelevant non-speech information. To address this, the paper proposes Segmental Average Pooling (SAP), which selectively focuses on informative speech segments while ignoring non-speech segments. By applying both GAP and SAP to SSL features, our approach utilizes overall speech signal information from GAP and specific information from SAP, leading to improved SER performance. Experiments show state-of-the-art results on the IEMOCAP for English and superior performance on KEMDy19 for Korean datasets in both unweighted and weighted accuracies.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12416
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
Hyeon, Jonghwan
Oh, Yung-Hwan
Choi, Ho-Jin
Sound
Artificial Intelligence
Audio and Speech Processing
Speech Emotion Recognition (SER) analyzes human emotions expressed through speech. Self-supervised learning (SSL) offers a promising approach to SER by learning meaningful representations from a large amount of unlabeled audio data. However, existing SSL-based methods rely on Global Average Pooling (GAP) to represent audio signals, treating speech and non-speech segments equally. This can lead to dilution of informative speech features by irrelevant non-speech information. To address this, the paper proposes Segmental Average Pooling (SAP), which selectively focuses on informative speech segments while ignoring non-speech segments. By applying both GAP and SAP to SSL features, our approach utilizes overall speech signal information from GAP and specific information from SAP, leading to improved SER performance. Experiments show state-of-the-art results on the IEMOCAP for English and superior performance on KEMDy19 for Korean datasets in both unweighted and weighted accuracies.
title Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2410.12416