DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Alexander H., Chang, Heng-Jui, Auli, Michael, Hsu, Wei-Ning, Glass, James R.
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911757561233408
author Liu, Alexander H.
Chang, Heng-Jui
Auli, Michael
Hsu, Wei-Ning
Glass, James R.
author_facet Liu, Alexander H.
Chang, Heng-Jui
Auli, Michael
Hsu, Wei-Ning
Glass, James R.
contents In this paper, we introduce self-distillation and online clustering for self-supervised speech representation learning (DinoSR) which combines masked language modeling, self-distillation, and online clustering. We show that these concepts complement each other and result in a strong representation learning model for speech. DinoSR first extracts contextualized embeddings from the input audio with a teacher network, then runs an online clustering system on the embeddings to yield a machine-discovered phone inventory, and finally uses the discretized tokens to guide a student network. We show that DinoSR surpasses previous state-of-the-art performance in several downstream tasks, and provide a detailed analysis of the model and the learned discrete units.
format Preprint
id arxiv_https___arxiv_org_abs_2305_10005
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning
Liu, Alexander H.
Chang, Heng-Jui
Auli, Michael
Hsu, Wei-Ning
Glass, James R.
Computation and Language
In this paper, we introduce self-distillation and online clustering for self-supervised speech representation learning (DinoSR) which combines masked language modeling, self-distillation, and online clustering. We show that these concepts complement each other and result in a strong representation learning model for speech. DinoSR first extracts contextualized embeddings from the input audio with a teacher network, then runs an online clustering system on the embeddings to yield a machine-discovered phone inventory, and finally uses the discretized tokens to guide a student network. We show that DinoSR surpasses previous state-of-the-art performance in several downstream tasks, and provide a detailed analysis of the model and the learned discrete units.
title DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning
topic Computation and Language
url https://arxiv.org/abs/2305.10005