Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Azzouz, Sofiane, Vuissoz, Pierre-André, Laprie, Yves
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910050523545600
author Azzouz, Sofiane
Vuissoz, Pierre-André
Laprie, Yves
author_facet Azzouz, Sofiane
Vuissoz, Pierre-André
Laprie, Yves
contents Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a comparison to the baseline introduced in our previous work, which is based on Mel-Frequency Cepstral Coefficients (MFCCs). All the approaches considered are based on a denoised speech signal and aim to investigate the impact of incorporating phonetic information through three successive levels: an uncorrected automatic transcription, a temporally aligned phonetic segmentation, and an expert manual correction following alignment. The models are trained to predict articulatory contours extracted from vocal tract MRI images using an automatic contour tracking method. The results show that, among the models relying on phonetic representations, manual correction after alignment yields the best performance, approaching that of the baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2603_11847
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data
Azzouz, Sofiane
Vuissoz, Pierre-André
Laprie, Yves
Audio and Speech Processing
Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a comparison to the baseline introduced in our previous work, which is based on Mel-Frequency Cepstral Coefficients (MFCCs). All the approaches considered are based on a denoised speech signal and aim to investigate the impact of incorporating phonetic information through three successive levels: an uncorrected automatic transcription, a temporally aligned phonetic segmentation, and an expert manual correction following alignment. The models are trained to predict articulatory contours extracted from vocal tract MRI images using an automatic contour tracking method. The results show that, among the models relying on phonetic representations, manual correction after alignment yields the best performance, approaching that of the baseline.
title Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data
topic Audio and Speech Processing
url https://arxiv.org/abs/2603.11847