Conformer-based Ultrasound-to-Speech Conversion

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ibrahimov, Ibrahim, Csaba, Zainkó, Gosztolya, Gábor
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918045462560768
author Ibrahimov, Ibrahim
Csaba, Zainkó
Gosztolya, Gábor
author_facet Ibrahimov, Ibrahim
Csaba, Zainkó
Gosztolya, Gábor
contents Deep neural networks have shown promising potential for ultrasound-to-speech conversion task towards Silent Speech Interfaces. In this work, we applied two Conformer-based DNN architectures (Base and one with bi-LSTM) for this task. Speaker-specific models were trained on the data of four speakers from the Ultrasuite-Tal80 dataset, while the generated mel spectrograms were synthesized to audio waveform using a HiFi-GAN vocoder. Compared to a standard 2D-CNN baseline, objective measurements (MSE and mel cepstral distortion) showed no statistically significant improvement for either model. However, a MUSHRA listening test revealed that Conformer with bi-LSTM provided better perceptual quality, while Conformer Base matched the performance of the baseline along with a 3x faster training time due to its simpler architecture. These findings suggest that Conformer-based models, especially the Conformer with bi-LSTM, offer a promising alternative to CNNs for ultrasound-to-speech conversion.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03831
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Conformer-based Ultrasound-to-Speech Conversion
Ibrahimov, Ibrahim
Csaba, Zainkó
Gosztolya, Gábor
Sound
Multimedia
Audio and Speech Processing
Deep neural networks have shown promising potential for ultrasound-to-speech conversion task towards Silent Speech Interfaces. In this work, we applied two Conformer-based DNN architectures (Base and one with bi-LSTM) for this task. Speaker-specific models were trained on the data of four speakers from the Ultrasuite-Tal80 dataset, while the generated mel spectrograms were synthesized to audio waveform using a HiFi-GAN vocoder. Compared to a standard 2D-CNN baseline, objective measurements (MSE and mel cepstral distortion) showed no statistically significant improvement for either model. However, a MUSHRA listening test revealed that Conformer with bi-LSTM provided better perceptual quality, while Conformer Base matched the performance of the baseline along with a 3x faster training time due to its simpler architecture. These findings suggest that Conformer-based models, especially the Conformer with bi-LSTM, offer a promising alternative to CNNs for ultrasound-to-speech conversion.
title Conformer-based Ultrasound-to-Speech Conversion
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2506.03831