Language Bias in Self-Supervised Learning For Automatic Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Storey, Edward, Harte, Naomi, Bell, Peter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910807461199872
author Storey, Edward
Harte, Naomi
Bell, Peter
author_facet Storey, Edward
Harte, Naomi
Bell, Peter
contents Self-supervised learning (SSL) is used in deep learning to train on large datasets without the need for expensive labelling of the data. Recently, large Automatic Speech Recognition (ASR) models such as XLS-R have utilised SSL to train on over one hundred different languages simultaneously. However, deeper investigation shows that the bulk of the training data for XLS-R comes from a small number of languages. Biases learned through SSL have been shown to exist in multiple domains, but language bias in multilingual SSL ASR has not been thoroughly examined. In this paper, we utilise the Lottery Ticket Hypothesis (LTH) to identify language-specific subnetworks within XLS-R and test the performance of these subnetworks on a variety of different languages. We are able to show that when fine-tuning, XLS-R bypasses traditional linguistic knowledge and builds only on weights learned from the languages with the largest data contribution to the pretraining data.
format Preprint
id arxiv_https___arxiv_org_abs_2501_19321
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language Bias in Self-Supervised Learning For Automatic Speech Recognition
Storey, Edward
Harte, Naomi
Bell, Peter
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Signal Processing
Self-supervised learning (SSL) is used in deep learning to train on large datasets without the need for expensive labelling of the data. Recently, large Automatic Speech Recognition (ASR) models such as XLS-R have utilised SSL to train on over one hundred different languages simultaneously. However, deeper investigation shows that the bulk of the training data for XLS-R comes from a small number of languages. Biases learned through SSL have been shown to exist in multiple domains, but language bias in multilingual SSL ASR has not been thoroughly examined. In this paper, we utilise the Lottery Ticket Hypothesis (LTH) to identify language-specific subnetworks within XLS-R and test the performance of these subnetworks on a variety of different languages. We are able to show that when fine-tuning, XLS-R bypasses traditional linguistic knowledge and builds only on weights learned from the languages with the largest data contribution to the pretraining data.
title Language Bias in Self-Supervised Learning For Automatic Speech Recognition
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Machine Learning
Signal Processing
url https://arxiv.org/abs/2501.19321