CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tsai, Yun-Shao, Lin, Yi-Cheng, Chou, Huang-Cheng, Lee, Hung-yi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911265300938752
author Tsai, Yun-Shao
Lin, Yi-Cheng
Chou, Huang-Cheng
Lee, Hung-yi
author_facet Tsai, Yun-Shao
Lin, Yi-Cheng
Chou, Huang-Cheng
Lee, Hung-yi
contents Bias in speech emotion recognition (SER) systems often stems from spurious correlations between speaker characteristics and emotional labels, leading to unfair predictions across demographic groups. Many existing debiasing methods require model-specific changes or demographic annotations, limiting their practical use. We present CO-VADA, a Confidence-Oriented Voice Augmentation Debiasing Approach that mitigates bias without modifying model architecture or relying on demographic information. CO-VADA identifies training samples that reflect bias patterns present in the training data and then applies voice conversion to alter irrelevant attributes and generate samples. These augmented samples introduce speaker variations that differ from dominant patterns in the data, guiding the model to focus more on emotion-relevant features. Our framework is compatible with various SER models and voice conversion tools, making it a scalable and practical solution for improving fairness in SER systems.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06071
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
Tsai, Yun-Shao
Lin, Yi-Cheng
Chou, Huang-Cheng
Lee, Hung-yi
Audio and Speech Processing
Computation and Language
Bias in speech emotion recognition (SER) systems often stems from spurious correlations between speaker characteristics and emotional labels, leading to unfair predictions across demographic groups. Many existing debiasing methods require model-specific changes or demographic annotations, limiting their practical use. We present CO-VADA, a Confidence-Oriented Voice Augmentation Debiasing Approach that mitigates bias without modifying model architecture or relying on demographic information. CO-VADA identifies training samples that reflect bias patterns present in the training data and then applies voice conversion to alter irrelevant attributes and generate samples. These augmented samples introduce speaker variations that differ from dominant patterns in the data, guiding the model to focus more on emotion-relevant features. Our framework is compatible with various SER models and voice conversion tools, making it a scalable and practical solution for improving fairness in SER systems.
title CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2506.06071