ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dutta, Soumya, Balaji, Smruthi, R, Varada, Salinamakki, Viveka, Ganapathy, Sriram
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912391918256128
author Dutta, Soumya
Balaji, Smruthi
R, Varada
Salinamakki, Viveka
Ganapathy, Sriram
author_facet Dutta, Soumya
Balaji, Smruthi
R, Varada
Salinamakki, Viveka
Ganapathy, Sriram
contents Speech emotion recognition (SER) in naturalistic settings remains a challenge due to the intrinsic variability, diverse recording conditions, and class imbalance. As participants in the Interspeech Naturalistic SER Challenge which focused on these complexities, we present Abhinaya, a system integrating speech-based, text-based, and speech-text models. Our approach fine-tunes self-supervised and speech large language models (SLLM) for speech representations, leverages large language models (LLM) for textual context, and employs speech-text modeling with an SLLM to capture nuanced emotional cues. To combat class imbalance, we apply tailored loss functions and generate categorical decisions through majority voting. Despite one model not being fully trained, the Abhinaya system ranked 4th among 166 submissions. Upon completion of training, it achieved state-of-the-art performance among published results, demonstrating the effectiveness of our approach for SER in real-world conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18217
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
Dutta, Soumya
Balaji, Smruthi
R, Varada
Salinamakki, Viveka
Ganapathy, Sriram
Sound
Artificial Intelligence
Audio and Speech Processing
Speech emotion recognition (SER) in naturalistic settings remains a challenge due to the intrinsic variability, diverse recording conditions, and class imbalance. As participants in the Interspeech Naturalistic SER Challenge which focused on these complexities, we present Abhinaya, a system integrating speech-based, text-based, and speech-text models. Our approach fine-tunes self-supervised and speech large language models (SLLM) for speech representations, leverages large language models (LLM) for textual context, and employs speech-text modeling with an SLLM to capture nuanced emotional cues. To combat class imbalance, we apply tailored loss functions and generate categorical decisions through majority voting. Despite one model not being fully trained, the Abhinaya system ranked 4th among 166 submissions. Upon completion of training, it achieved state-of-the-art performance among published results, demonstrating the effectiveness of our approach for SER in real-world conditions.
title ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.18217