Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Weise, Tobias, Klumpp, Philipp, Demir, Kubilay Can, Pérez-Toro, Paula Andrea, Schuster, Maria, Noeth, Elmar, Heismann, Bjoern, Maier, Andreas, Yang, Seung Hee
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910511690416128
author Weise, Tobias
Klumpp, Philipp
Demir, Kubilay Can
Pérez-Toro, Paula Andrea
Schuster, Maria
Noeth, Elmar
Heismann, Bjoern
Maier, Andreas
Yang, Seung Hee
author_facet Weise, Tobias
Klumpp, Philipp
Demir, Kubilay Can
Pérez-Toro, Paula Andrea
Schuster, Maria
Noeth, Elmar
Heismann, Bjoern
Maier, Andreas
Yang, Seung Hee
contents This paper introduces a novel combination of two tasks, previously treated separately: acoustic-to-articulatory speech inversion (AAI) and phoneme-to-articulatory (PTA) motion estimation. We refer to this joint task as acoustic phoneme-to-articulatory speech inversion (APTAI) and explore two different approaches, both working speaker- and text-independently during inference. We use a multi-task learning setup, with the end-to-end goal of taking raw speech as input and estimating the corresponding articulatory movements, phoneme sequence, and phoneme alignment. While both proposed approaches share these same requirements, they differ in their way of achieving phoneme-related predictions: one is based on frame classification, the other on a two-staged training procedure and forced alignment. We reach competitive performance of 0.73 mean correlation for the AAI task and achieve up to approximately 87% frame overlap compared to a state-of-the-art text-dependent phoneme force aligner.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03132
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
Weise, Tobias
Klumpp, Philipp
Demir, Kubilay Can
Pérez-Toro, Paula Andrea
Schuster, Maria
Noeth, Elmar
Heismann, Bjoern
Maier, Andreas
Yang, Seung Hee
Sound
Artificial Intelligence
Computation and Language
Machine Learning
Audio and Speech Processing
This paper introduces a novel combination of two tasks, previously treated separately: acoustic-to-articulatory speech inversion (AAI) and phoneme-to-articulatory (PTA) motion estimation. We refer to this joint task as acoustic phoneme-to-articulatory speech inversion (APTAI) and explore two different approaches, both working speaker- and text-independently during inference. We use a multi-task learning setup, with the end-to-end goal of taking raw speech as input and estimating the corresponding articulatory movements, phoneme sequence, and phoneme alignment. While both proposed approaches share these same requirements, they differ in their way of achieving phoneme-related predictions: one is based on frame classification, the other on a two-staged training procedure and forced alignment. We reach competitive performance of 0.73 mean correlation for the AAI task and achieve up to approximately 87% frame overlap compared to a state-of-the-art text-dependent phoneme force aligner.
title Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
topic Sound
Artificial Intelligence
Computation and Language
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2407.03132