Self-Supervised Models of Speech Infer Universal Articulatory Kinematics

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cho, Cheol Jun, Mohamed, Abdelrahman, Black, Alan W, Anumanchipalli, Gopala K.
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914641257431040
author Cho, Cheol Jun
Mohamed, Abdelrahman
Black, Alan W
Anumanchipalli, Gopala K.
author_facet Cho, Cheol Jun
Mohamed, Abdelrahman
Black, Alan W
Anumanchipalli, Gopala K.
contents Self-Supervised Learning (SSL) based models of speech have shown remarkable performance on a range of downstream tasks. These state-of-the-art models have remained blackboxes, but many recent studies have begun "probing" models like HuBERT, to correlate their internal representations to different aspects of speech. In this paper, we show "inference of articulatory kinematics" as fundamental property of SSL models, i.e., the ability of these models to transform acoustics into the causal articulatory dynamics underlying the speech signal. We also show that this abstraction is largely overlapping across the language of the data used to train the model, with preference to the language with similar phonological system. Furthermore, we show that with simple affine transformations, Acoustic-to-Articulatory inversion (AAI) is transferrable across speakers, even across genders, languages, and dialects, showing the generalizability of this property. Together, these results shed new light on the internals of SSL models that are critical to their superior performance, and open up new avenues into language-agnostic universal models for speech engineering, that are interpretable and grounded in speech science.
format Preprint
id arxiv_https___arxiv_org_abs_2310_10788
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
Cho, Cheol Jun
Mohamed, Abdelrahman
Black, Alan W
Anumanchipalli, Gopala K.
Audio and Speech Processing
Computation and Language
Self-Supervised Learning (SSL) based models of speech have shown remarkable performance on a range of downstream tasks. These state-of-the-art models have remained blackboxes, but many recent studies have begun "probing" models like HuBERT, to correlate their internal representations to different aspects of speech. In this paper, we show "inference of articulatory kinematics" as fundamental property of SSL models, i.e., the ability of these models to transform acoustics into the causal articulatory dynamics underlying the speech signal. We also show that this abstraction is largely overlapping across the language of the data used to train the model, with preference to the language with similar phonological system. Furthermore, we show that with simple affine transformations, Acoustic-to-Articulatory inversion (AAI) is transferrable across speakers, even across genders, languages, and dialects, showing the generalizability of this property. Together, these results shed new light on the internals of SSL models that are critical to their superior performance, and open up new avenues into language-agnostic universal models for speech engineering, that are interpretable and grounded in speech science.
title Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
topic Audio and Speech Processing
Computation and Language
url https://arxiv.org/abs/2310.10788