Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tzeng, Jing-Tong, Su, Bo-Hao, Wu, Ya-Tse, Chou, Hsing-Hang, Lee, Chi-Chun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916969040576512
author Tzeng, Jing-Tong
Su, Bo-Hao
Wu, Ya-Tse
Chou, Hsing-Hang
Lee, Chi-Chun
author_facet Tzeng, Jing-Tong
Su, Bo-Hao
Wu, Ya-Tse
Chou, Hsing-Hang
Lee, Chi-Chun
contents In this study, we revisit key training strategies in machine learning often overlooked in favor of deeper architectures. Specifically, we explore balancing strategies, activation functions, and fine-tuning techniques to enhance speech emotion recognition (SER) in naturalistic conditions. Our findings show that simple modifications improve generalization with minimal architectural changes. Our multi-modal fusion model, integrating these optimizations, achieves a valence CCC of 0.6953, the best valence score in Task 2: Emotional Attribute Regression. Notably, fine-tuning RoBERTa and WavLM separately in a single-modality setting, followed by feature fusion without training the backbone extractor, yields the highest valence performance. Additionally, focal loss and activation functions significantly enhance performance without increasing complexity. These results suggest that refining core components, rather than deepening models, leads to more robust SER in-the-wild.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07282
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild
Tzeng, Jing-Tong
Su, Bo-Hao
Wu, Ya-Tse
Chou, Hsing-Hang
Lee, Chi-Chun
Audio and Speech Processing
In this study, we revisit key training strategies in machine learning often overlooked in favor of deeper architectures. Specifically, we explore balancing strategies, activation functions, and fine-tuning techniques to enhance speech emotion recognition (SER) in naturalistic conditions. Our findings show that simple modifications improve generalization with minimal architectural changes. Our multi-modal fusion model, integrating these optimizations, achieves a valence CCC of 0.6953, the best valence score in Task 2: Emotional Attribute Regression. Notably, fine-tuning RoBERTa and WavLM separately in a single-modality setting, followed by feature fusion without training the backbone extractor, yields the highest valence performance. Additionally, focal loss and activation functions significantly enhance performance without increasing complexity. These results suggest that refining core components, rather than deepening models, leads to more robust SER in-the-wild.
title Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.07282