EmoHRNet: High-Resolution Neural Network Based Speech Emotion Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Muppidi, Akshay, Radfar, Martin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915536704634880
author Muppidi, Akshay
Radfar, Martin
author_facet Muppidi, Akshay
Radfar, Martin
contents Speech emotion recognition (SER) is pivotal for enhancing human-machine interactions. This paper introduces "EmoHRNet", a novel adaptation of High-Resolution Networks (HRNet) tailored for SER. The HRNet structure is designed to maintain high-resolution representations from the initial to the final layers. By transforming audio samples into spectrograms, EmoHRNet leverages the HRNet architecture to extract high-level features. EmoHRNet's unique architecture maintains high-resolution representations throughout, capturing both granular and overarching emotional cues from speech signals. The model outperforms leading models, achieving accuracies of 92.45% on RAVDESS, 80.06% on IEMOCAP, and 92.77% on EMOVO. Thus, we show that EmoHRNet sets a new benchmark in the SER domain.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06072
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EmoHRNet: High-Resolution Neural Network Based Speech Emotion Recognition
Muppidi, Akshay
Radfar, Martin
Sound
Machine Learning
Speech emotion recognition (SER) is pivotal for enhancing human-machine interactions. This paper introduces "EmoHRNet", a novel adaptation of High-Resolution Networks (HRNet) tailored for SER. The HRNet structure is designed to maintain high-resolution representations from the initial to the final layers. By transforming audio samples into spectrograms, EmoHRNet leverages the HRNet architecture to extract high-level features. EmoHRNet's unique architecture maintains high-resolution representations throughout, capturing both granular and overarching emotional cues from speech signals. The model outperforms leading models, achieving accuracies of 92.45% on RAVDESS, 80.06% on IEMOCAP, and 92.77% on EMOVO. Thus, we show that EmoHRNet sets a new benchmark in the SER domain.
title EmoHRNet: High-Resolution Neural Network Based Speech Emotion Recognition
topic Sound
Machine Learning
url https://arxiv.org/abs/2510.06072