A Speech-to-Video Synthesis Approach Using Spatio-Temporal Diffusion for Vocal Tract MRI
Fuente:
arXiv
Saved in:
| Main Authors: | Pérez-Toro, Paula Andrea, Arias-Vergara, Tomás, Xing, Fangxu, Liu, Xiaofeng, Stone, Maureen, Zhuo, Jiachen, Orozco-Arroyave, Juan Rafael, Nöth, Elmar, Hutter, Jana, Prince, Jerry L., Maier, Andreas, Woo, Jonghye |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VocSegMRI: Multimodal Learning for Precise Vocal Tract Segmentation in Real-time MRI
by: Liu, Daiqi, et al.
Published: (2025)
by: Liu, Daiqi, et al.
Published: (2025)
Speech motion anomaly detection via cross-modal translation of 4D motion fields from tagged MRI
by: Liu, Xiaofeng, et al.
Published: (2024)
by: Liu, Xiaofeng, et al.
Published: (2024)
Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI
by: Liu, Daiqi, et al.
Published: (2026)
by: Liu, Daiqi, et al.
Published: (2026)
SIREM: Speech-Informed MRI Reconstruction with Learned Sampling
by: Hasan, Md, et al.
Published: (2026)
by: Hasan, Md, et al.
Published: (2026)
Speech Audio Generation from dynamic MRI via a Knowledge Enhanced Conditional Variational Autoencoder
by: Li, Yaxuan, et al.
Published: (2025)
by: Li, Yaxuan, et al.
Published: (2025)
Brightness-Invariant Tracking Estimation in Tagged MRI
by: Bian, Zhangxing, et al.
Published: (2025)
by: Bian, Zhangxing, et al.
Published: (2025)
Anatomy-based quality metric of diffusion-weighted MRI data for accurate derivation of muscle fiber orientation
by: Shusharina, Nadya, et al.
Published: (2024)
by: Shusharina, Nadya, et al.
Published: (2024)
Semi-Supervised Bone Marrow Lesion Detection from Knee MRI Segmentation Using Mask Inpainting Models
by: Qin, Shihua, et al.
Published: (2024)
by: Qin, Shihua, et al.
Published: (2024)
Audio-Vision Contrastive Learning for Phonological Class Recognition
by: Liu, Daiqi, et al.
Published: (2025)
by: Liu, Daiqi, et al.
Published: (2025)
Bias and Fairness in Self-Supervised Acoustic Representations for Cognitive Impairment Detection
by: Gulzar, Kashaf, et al.
Published: (2026)
by: Gulzar, Kashaf, et al.
Published: (2026)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
by: Nguyen, Hong, et al.
Published: (2024)
by: Nguyen, Hong, et al.
Published: (2024)
Principled Feature Disentanglement for High-Fidelity Unified Brain MRI Synthesis
by: Cho, Jihoon, et al.
Published: (2024)
by: Cho, Jihoon, et al.
Published: (2024)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
by: Hernandez, Abner, et al.
Published: (2026)
by: Hernandez, Abner, et al.
Published: (2026)
Reconstruction of the Vocal Tract from Speech via Phonetic Representations Using MRI Data
by: Azzouz, Sofiane, et al.
Published: (2026)
by: Azzouz, Sofiane, et al.
Published: (2026)
Coding Speech through Vocal Tract Kinematics
by: Cho, Cheol Jun, et al.
Published: (2024)
by: Cho, Cheol Jun, et al.
Published: (2024)
Disentangled Multimodal Brain MR Image Translation via Transformer-based Modality Infuser
by: Cho, Jihoon, et al.
Published: (2024)
by: Cho, Jihoon, et al.
Published: (2024)
Is Registering Raw Tagged-MR Enough for Strain Estimation in the Era of Deep Learning?
by: Bian, Zhangxing, et al.
Published: (2024)
by: Bian, Zhangxing, et al.
Published: (2024)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
by: Li, Chin-Jou, et al.
Published: (2025)
by: Li, Chin-Jou, et al.
Published: (2025)
Confidence-Guided Error Correction for Disordered Speech Recognition
by: Hernandez, Abner, et al.
Published: (2025)
by: Hernandez, Abner, et al.
Published: (2025)
RATNUS: Rapid, Automatic Thalamic Nuclei Segmentation using Multimodal MRI inputs
by: Feng, Anqi, et al.
Published: (2024)
by: Feng, Anqi, et al.
Published: (2024)
Ethics of Generating Synthetic MRI Vocal Tract Views from the Face
by: Shahid, Muhammad Suhaib, et al.
Published: (2024)
by: Shahid, Muhammad Suhaib, et al.
Published: (2024)
The Impact of Speech Anonymization on Pathology and Its Limits
by: Arasteh, Soroosh Tayebi, et al.
Published: (2024)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2024)
Multilingual Phonological Feature Recognition with Self-Supervised Speech Models
by: Hernandez, Abner, et al.
Published: (2026)
by: Hernandez, Abner, et al.
Published: (2026)
Large Language Models for Dysfluency Detection in Stuttered Speech
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
Treatment-wise Glioblastoma Survival Inference with Multi-parametric Preoperative MRI
by: Liu, Xiaofeng, et al.
Published: (2024)
by: Liu, Xiaofeng, et al.
Published: (2024)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
by: Wagner, Dominik, et al.
Published: (2025)
by: Wagner, Dominik, et al.
Published: (2025)
Multimodal Segmentation for Vocal Tract Modeling
by: Jain, Rishi, et al.
Published: (2024)
by: Jain, Rishi, et al.
Published: (2024)
Point-supervised Brain Tumor Segmentation with Box-prompted MedSAM
by: Liu, Xiaofeng, et al.
Published: (2024)
by: Liu, Xiaofeng, et al.
Published: (2024)
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
by: Bayerl, Sebastian P., et al.
Published: (2022)
by: Bayerl, Sebastian P., et al.
Published: (2022)
Label-Efficient 3D Brain Segmentation via Complementary 2D Diffusion Models with Orthogonal Views
by: Cho, Jihoon, et al.
Published: (2024)
by: Cho, Jihoon, et al.
Published: (2024)
Machine Learning Detection of Scarring Events in Killer Whales
by: Alexander Barnhill, et al.
Published: (2026)
by: Alexander Barnhill, et al.
Published: (2026)
Auditory Representation Effective for Estimating Vocal Tract Information
by: Irino, Toshio, et al.
Published: (2023)
by: Irino, Toshio, et al.
Published: (2023)
CATNUS: Coordinate-Aware Thalamic Nuclei Segmentation Using T1-Weighted MRI
by: Feng, Anqi, et al.
Published: (2025)
by: Feng, Anqi, et al.
Published: (2025)
Acoustic-to-articulatory Inversion of the Complete Vocal Tract from RT-MRI with Various Audio Embeddings and Dataset Sizes
by: Azzouz, Sofiane, et al.
Published: (2026)
by: Azzouz, Sofiane, et al.
Published: (2026)
Reconstruction of the Complete Vocal Tract Contour Through Acoustic to Articulatory Inversion Using Real-Time MRI Data
by: Azzouz, Sofiane, et al.
Published: (2025)
by: Azzouz, Sofiane, et al.
Published: (2025)
O fundamento estrutural do pensamento de Umberto Eco
by: Winfried Nöth
Published: (2016)
by: Winfried Nöth
Published: (2016)
Os discursos literários, científicos e filosóficos em C. S. Peirce1
by: Winfried Nöth
Published: (2023)
by: Winfried Nöth
Published: (2023)
AUTORREFERENCIALIDAD EN LA CRISIS DE LA MODERNIDAD
by: Winfried Nöth
Published: (2001)
by: Winfried Nöth
Published: (2001)
Comunicação: os paradigmas da simetria, antissimetria e assimetria
by: Winfried Nöth
Published: (2011)
by: Winfried Nöth
Published: (2011)
A teoria da comunicação de Charles S. Peirce e os equívocos de Ciro Marcondes Filho
by: Winfried Nöth
Published: (2013)
by: Winfried Nöth
Published: (2013)
Similar Items
-
VocSegMRI: Multimodal Learning for Precise Vocal Tract Segmentation in Real-time MRI
by: Liu, Daiqi, et al.
Published: (2025) -
Speech motion anomaly detection via cross-modal translation of 4D motion fields from tagged MRI
by: Liu, Xiaofeng, et al.
Published: (2024) -
Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI
by: Liu, Daiqi, et al.
Published: (2026) -
SIREM: Speech-Informed MRI Reconstruction with Learned Sampling
by: Hasan, Md, et al.
Published: (2026) -
Speech Audio Generation from dynamic MRI via a Knowledge Enhanced Conditional Variational Autoencoder
by: Li, Yaxuan, et al.
Published: (2025)