VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kumar, Puneet, Malik, Sarthak, Raman, Balasubramanian, Li, Xiaobai
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909621413740544
author Kumar, Puneet
Malik, Sarthak
Raman, Balasubramanian
Li, Xiaobai
author_facet Kumar, Puneet
Malik, Sarthak
Raman, Balasubramanian
Li, Xiaobai
contents This paper proposes a multimodal emotion recognition system, VIsual Spoken Textual Additive Net (VISTANet), to classify emotions reflected by input containing image, speech, and text into discrete classes. A new interpretability technique, K-Average Additive exPlanation (KAAP), has been developed that identifies important visual, spoken, and textual features leading to predicting a particular emotion class. The VISTANet fuses information from image, speech, and text modalities using a hybrid of intermediate and late fusion. It automatically adjusts the weights of their intermediate outputs while computing the weighted average. The KAAP technique computes the contribution of each modality and corresponding features toward predicting a particular emotion class. To mitigate the insufficiency of multimodal emotion datasets labelled with discrete emotion classes, we have constructed the IIT-R MMEmoRec dataset consisting of images, corresponding speech and text, and emotion labels ('angry,' 'happy,' 'hate,' and 'sad'). The VISTANet has resulted in an overall emotion recognition accuracy of 80.11% on the IIT-R MMEmoRec dataset using visual, spoken, and textual modalities, outperforming single or dual-modality configurations. The code and data can be accessed at https://github.com/MIntelligence-Group/MMEmoRec.
format Preprint
id arxiv_https___arxiv_org_abs_2208_11450
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition
Kumar, Puneet
Malik, Sarthak
Raman, Balasubramanian
Li, Xiaobai
Computer Vision and Pattern Recognition
This paper proposes a multimodal emotion recognition system, VIsual Spoken Textual Additive Net (VISTANet), to classify emotions reflected by input containing image, speech, and text into discrete classes. A new interpretability technique, K-Average Additive exPlanation (KAAP), has been developed that identifies important visual, spoken, and textual features leading to predicting a particular emotion class. The VISTANet fuses information from image, speech, and text modalities using a hybrid of intermediate and late fusion. It automatically adjusts the weights of their intermediate outputs while computing the weighted average. The KAAP technique computes the contribution of each modality and corresponding features toward predicting a particular emotion class. To mitigate the insufficiency of multimodal emotion datasets labelled with discrete emotion classes, we have constructed the IIT-R MMEmoRec dataset consisting of images, corresponding speech and text, and emotion labels ('angry,' 'happy,' 'hate,' and 'sad'). The VISTANet has resulted in an overall emotion recognition accuracy of 80.11% on the IIT-R MMEmoRec dataset using visual, spoken, and textual modalities, outperforming single or dual-modality configurations. The code and data can be accessed at https://github.com/MIntelligence-Group/MMEmoRec.
title VISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2208.11450