End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yamashita, Natsuo, Yamamoto, Masaaki, Kawaguchi, Yohei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917806223654912
author Yamashita, Natsuo
Yamamoto, Masaaki
Kawaguchi, Yohei
author_facet Yamashita, Natsuo
Yamamoto, Masaaki
Kawaguchi, Yohei
contents Speech Emotion Recognition (SER) often operates on speech segments detected by a Voice Activity Detection (VAD) model. However, VAD models may output flawed speech segments, especially in noisy environments, resulting in degraded performance of subsequent SER models. To address this issue, we propose an end-to-end (E2E) method that integrates VAD and SER using Self-Supervised Learning (SSL) features. The VAD module first receives the SSL features as input, and the segmented SSL features are then fed into the SER module. Both the VAD and SER modules are jointly trained to optimize SER performance. Experimental results on the IEMOCAP dataset demonstrate that our proposed method improves SER performance. Furthermore, to investigate the effect of our proposed method on the VAD and SSL modules, we present an analysis of the VAD outputs and the weights of each layer of the SSL encoder.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13282
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
Yamashita, Natsuo
Yamamoto, Masaaki
Kawaguchi, Yohei
Sound
Audio and Speech Processing
Speech Emotion Recognition (SER) often operates on speech segments detected by a Voice Activity Detection (VAD) model. However, VAD models may output flawed speech segments, especially in noisy environments, resulting in degraded performance of subsequent SER models. To address this issue, we propose an end-to-end (E2E) method that integrates VAD and SER using Self-Supervised Learning (SSL) features. The VAD module first receives the SSL features as input, and the segmented SSL features are then fed into the SER module. Both the VAD and SER modules are jointly trained to optimize SER performance. Experimental results on the IEMOCAP dataset demonstrate that our proposed method improves SER performance. Furthermore, to investigate the effect of our proposed method on the VAD and SSL modules, we present an analysis of the VAD outputs and the weights of each layer of the SSL encoder.
title End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.13282