Multi-Microphone and Multi-Modal Emotion Recognition in Reverberant Environment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cohen, Ohad, Hazan, Gershon, Gannot, Sharon
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909709600030720
author Cohen, Ohad
Hazan, Gershon
Gannot, Sharon
author_facet Cohen, Ohad
Hazan, Gershon
Gannot, Sharon
contents This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio Transformer (HTS-AT) for multi-channel audio processing with an R(2+1)D Convolutional Neural Networks (CNN) model for video analysis. We evaluate our proposed method on a reverberated version of the Ryerson audio-visual database of emotional speech and song (RAVDESS) dataset using synthetic and real-world Room Impulse Responsess (RIRs). Our results demonstrate that integrating audio and video modalities yields superior performance compared to uni-modal approaches, especially in challenging acoustic conditions. Moreover, we show that the multimodal (audiovisual) approach that utilizes multiple microphones outperforms its single-microphone counterpart.
format Preprint
id arxiv_https___arxiv_org_abs_2409_09545
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Microphone and Multi-Modal Emotion Recognition in Reverberant Environment
Cohen, Ohad
Hazan, Gershon
Gannot, Sharon
Sound
Machine Learning
Audio and Speech Processing
This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio Transformer (HTS-AT) for multi-channel audio processing with an R(2+1)D Convolutional Neural Networks (CNN) model for video analysis. We evaluate our proposed method on a reverberated version of the Ryerson audio-visual database of emotional speech and song (RAVDESS) dataset using synthetic and real-world Room Impulse Responsess (RIRs). Our results demonstrate that integrating audio and video modalities yields superior performance compared to uni-modal approaches, especially in challenging acoustic conditions. Moreover, we show that the multimodal (audiovisual) approach that utilizes multiple microphones outperforms its single-microphone counterpart.
title Multi-Microphone and Multi-Modal Emotion Recognition in Reverberant Environment
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2409.09545