Enrolment-based personalisation for improving individual-level fairness in speech emotion recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Triantafyllopoulos, Andreas, Schuller, Björn
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912318697242624
author Triantafyllopoulos, Andreas
Schuller, Björn
author_facet Triantafyllopoulos, Andreas
Schuller, Björn
contents The expression of emotion is highly individualistic. However, contemporary speech emotion recognition (SER) systems typically rely on population-level models that adopt a `one-size-fits-all' approach for predicting emotion. Moreover, standard evaluation practices measure performance also on the population level, thus failing to characterise how models work across different speakers. In the present contribution, we present a new method for capitalising on individual differences to adapt an SER model to each new speaker using a minimal set of enrolment utterances. In addition, we present novel evaluation schemes for measuring fairness across different speakers. Our findings show that aggregated evaluation metrics may obfuscate fairness issues on the individual-level, which are uncovered by our evaluation, and that our proposed method can improve performance both in aggregated and disaggregated terms.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06665
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enrolment-based personalisation for improving individual-level fairness in speech emotion recognition
Triantafyllopoulos, Andreas
Schuller, Björn
Computation and Language
The expression of emotion is highly individualistic. However, contemporary speech emotion recognition (SER) systems typically rely on population-level models that adopt a `one-size-fits-all' approach for predicting emotion. Moreover, standard evaluation practices measure performance also on the population level, thus failing to characterise how models work across different speakers. In the present contribution, we present a new method for capitalising on individual differences to adapt an SER model to each new speaker using a minimal set of enrolment utterances. In addition, we present novel evaluation schemes for measuring fairness across different speakers. Our findings show that aggregated evaluation metrics may obfuscate fairness issues on the individual-level, which are uncovered by our evaluation, and that our proposed method can improve performance both in aggregated and disaggregated terms.
title Enrolment-based personalisation for improving individual-level fairness in speech emotion recognition
topic Computation and Language
url https://arxiv.org/abs/2406.06665