Training Large ASR Encoders with Differential Privacy

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chauhan, Geeticka, Chien, Steve, Thakkar, Om, Thakurta, Abhradeep, Narayanan, Arun
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917781549613056
author Chauhan, Geeticka
Chien, Steve
Thakkar, Om
Thakurta, Abhradeep
Narayanan, Arun
author_facet Chauhan, Geeticka
Chien, Steve
Thakkar, Om
Thakurta, Abhradeep
Narayanan, Arun
contents Self-supervised learning (SSL) methods for large speech models have proven to be highly effective at ASR. With the interest in public deployment of large pre-trained models, there is a rising concern for unintended memorization and leakage of sensitive data points from the training data. In this paper, we apply differentially private (DP) pre-training to a SOTA Conformer-based encoder, and study its performance on a downstream ASR task assuming the fine-tuning data is public. This paper is the first to apply DP to SSL for ASR, investigating the DP noise tolerance of the BEST-RQ pre-training method. Notably, we introduce a novel variant of model pruning called gradient-based layer freezing that provides strong improvements in privacy-utility-compute trade-offs. Our approach yields a LibriSpeech test-clean/other WER (%) of 3.78/ 8.41 with ($10$, 1e^-9)-DP for extrapolation towards low dataset scales, and 2.81/ 5.89 with (10, 7.9e^-11)-DP for extrapolation towards high scales.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13953
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Training Large ASR Encoders with Differential Privacy
Chauhan, Geeticka
Chien, Steve
Thakkar, Om
Thakurta, Abhradeep
Narayanan, Arun
Sound
Cryptography and Security
Machine Learning
Audio and Speech Processing
Self-supervised learning (SSL) methods for large speech models have proven to be highly effective at ASR. With the interest in public deployment of large pre-trained models, there is a rising concern for unintended memorization and leakage of sensitive data points from the training data. In this paper, we apply differentially private (DP) pre-training to a SOTA Conformer-based encoder, and study its performance on a downstream ASR task assuming the fine-tuning data is public. This paper is the first to apply DP to SSL for ASR, investigating the DP noise tolerance of the BEST-RQ pre-training method. Notably, we introduce a novel variant of model pruning called gradient-based layer freezing that provides strong improvements in privacy-utility-compute trade-offs. Our approach yields a LibriSpeech test-clean/other WER (%) of 3.78/ 8.41 with ($10$, 1e^-9)-DP for extrapolation towards low dataset scales, and 2.81/ 5.89 with (10, 7.9e^-11)-DP for extrapolation towards high scales.
title Training Large ASR Encoders with Differential Privacy
topic Sound
Cryptography and Security
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2409.13953