Contrastive Regularization for Accent-Robust ASR

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Thai, Van-Phat, Dhruv, Aradhya, Pham, Duc-Thinh, Alam, Sameer
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913089119584256
author Thai, Van-Phat
Dhruv, Aradhya
Pham, Duc-Thinh
Alam, Sameer
author_facet Thai, Van-Phat
Dhruv, Aradhya
Pham, Duc-Thinh
Alam, Sameer
contents ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability. We investigate supervised contrastive learning (SupCon) as a lightweight, accent-invariant auxiliary objective for CTC fine-tuning. An utterance-level contrastive loss regularizes encoder representations without architectural modification or explicit accent supervision. Experiments on the L2-ARCTIC benchmark show consistent WER reductions across multiple pretrained encoders, with up to 25 -- 29\% relative reduction under unseen-accent evaluation. Analysis using within-transcript cosine dispersion indicates that SupCon promotes more compact and stable representation geometry under accent variability. Overall, SupCon provides an effective and model-agnostic regularization strategy for improving accent robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2605_03297
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Contrastive Regularization for Accent-Robust ASR
Thai, Van-Phat
Dhruv, Aradhya
Pham, Duc-Thinh
Alam, Sameer
Sound
Machine Learning
ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability. We investigate supervised contrastive learning (SupCon) as a lightweight, accent-invariant auxiliary objective for CTC fine-tuning. An utterance-level contrastive loss regularizes encoder representations without architectural modification or explicit accent supervision. Experiments on the L2-ARCTIC benchmark show consistent WER reductions across multiple pretrained encoders, with up to 25 -- 29\% relative reduction under unseen-accent evaluation. Analysis using within-transcript cosine dispersion indicates that SupCon promotes more compact and stable representation geometry under accent variability. Overall, SupCon provides an effective and model-agnostic regularization strategy for improving accent robustness.
title Contrastive Regularization for Accent-Robust ASR
topic Sound
Machine Learning
url https://arxiv.org/abs/2605.03297