On the social bias of speech self-supervised models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lin, Yi-Cheng, Lin, Tzu-Quan, Lin, Hsi-Che, Liu, Andy T., Lee, Hung-yi
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911351104864256
author Lin, Yi-Cheng
Lin, Tzu-Quan
Lin, Hsi-Che
Liu, Andy T.
Lee, Hung-yi
author_facet Lin, Yi-Cheng
Lin, Tzu-Quan
Lin, Hsi-Che
Liu, Andy T.
Lee, Hung-yi
contents Self-supervised learning (SSL) speech models have achieved remarkable performance in various tasks, yet the biased outcomes, especially affecting marginalized groups, raise significant concerns. Social bias refers to the phenomenon where algorithms potentially amplify disparate properties between social groups present in the data used for training. Bias in SSL models can perpetuate injustice by automating discriminatory patterns and reinforcing inequitable systems. This work reveals that prevalent SSL models inadvertently acquire biased associations. We probe how various factors, such as model architecture, size, and training methodologies, influence the propagation of social bias within these models. Finally, we explore the efficacy of debiasing SSL models through regularization techniques, specifically via model compression. Our findings reveal that employing techniques such as row-pruning and training wider, shallower models can effectively mitigate social bias within SSL model.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04997
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the social bias of speech self-supervised models
Lin, Yi-Cheng
Lin, Tzu-Quan
Lin, Hsi-Che
Liu, Andy T.
Lee, Hung-yi
Audio and Speech Processing
Machine Learning
Self-supervised learning (SSL) speech models have achieved remarkable performance in various tasks, yet the biased outcomes, especially affecting marginalized groups, raise significant concerns. Social bias refers to the phenomenon where algorithms potentially amplify disparate properties between social groups present in the data used for training. Bias in SSL models can perpetuate injustice by automating discriminatory patterns and reinforcing inequitable systems. This work reveals that prevalent SSL models inadvertently acquire biased associations. We probe how various factors, such as model architecture, size, and training methodologies, influence the propagation of social bias within these models. Finally, we explore the efficacy of debiasing SSL models through regularization techniques, specifically via model compression. Our findings reveal that employing techniques such as row-pruning and training wider, shallower models can effectively mitigate social bias within SSL model.
title On the social bias of speech self-supervised models
topic Audio and Speech Processing
Machine Learning
url https://arxiv.org/abs/2406.04997