Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
Fuente:
arXiv
Saved in:
| Main Authors: | Weise, Tobias, Klumpp, Philipp, Demir, Kubilay Can, Pérez-Toro, Paula Andrea, Schuster, Maria, Noeth, Elmar, Heismann, Bjoern, Maier, Andreas, Yang, Seung Hee |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis
by: Demir, Kubilay Can, et al.
Published: (2024)
by: Demir, Kubilay Can, et al.
Published: (2024)
The Impact of Speech Anonymization on Pathology and Its Limits
by: Arasteh, Soroosh Tayebi, et al.
Published: (2024)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2024)
Towards a Quantitative Analysis of Coarticulation with a Phoneme-to-Articulatory Model
by: Fan, Chaofei, et al.
Published: (2024)
by: Fan, Chaofei, et al.
Published: (2024)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
by: Fujita, Kenichi, et al.
Published: (2024)
by: Fujita, Kenichi, et al.
Published: (2024)
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
by: Yang, Dong, et al.
Published: (2025)
by: Yang, Dong, et al.
Published: (2025)
Speaker-Independent Acoustic-to-Articulatory Inversion through Multi-Channel Attention Discriminator
by: Chung, Woo-Jin, et al.
Published: (2024)
by: Chung, Woo-Jin, et al.
Published: (2024)
MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI
by: Shah, Neil, et al.
Published: (2024)
by: Shah, Neil, et al.
Published: (2024)
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
by: Park, Jay, et al.
Published: (2025)
by: Park, Jay, et al.
Published: (2025)
Acoustic to Articulatory Speech Inversion for Children with Velopharyngeal Insufficiency
by: Tabatabaee, Saba, et al.
Published: (2025)
by: Tabatabaee, Saba, et al.
Published: (2025)
Enhancing Acoustic-to-Articulatory Speech Inversion by Incorporating Nasality
by: Tabatabaee, Saba, et al.
Published: (2025)
by: Tabatabaee, Saba, et al.
Published: (2025)
Perceptual implications of automatic anonymization in pathological speech
by: Arasteh, Soroosh Tayebi, et al.
Published: (2025)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2025)
Deep Speech Synthesis from Multimodal Articulatory Representations
by: Wu, Peter, et al.
Published: (2024)
by: Wu, Peter, et al.
Published: (2024)
Perceptual Ratings Predict Speech Inversion Articulatory Kinematics in Childhood Speech Sound Disorders
by: Benway, Nina R., et al.
Published: (2025)
by: Benway, Nina R., et al.
Published: (2025)
Multi-Label Training for Text-Independent Speaker Identification
by: Xue, Yuqi
Published: (2022)
by: Xue, Yuqi
Published: (2022)
Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech
by: Kolani, Yakov, et al.
Published: (2025)
by: Kolani, Yakov, et al.
Published: (2025)
Acoustic-to-Articulatory Inversion of Clean Speech Using an MRI-Trained Model
by: Azzouz, Sofiane, et al.
Published: (2026)
by: Azzouz, Sofiane, et al.
Published: (2026)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
by: Park, Nohil, et al.
Published: (2024)
by: Park, Nohil, et al.
Published: (2024)
Optimizing Speech-Input Length for Speaker-Independent Depression Classification
by: Rutowski, Tomasz, et al.
Published: (2024)
by: Rutowski, Tomasz, et al.
Published: (2024)
Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
by: Cho, Cheol Jun, et al.
Published: (2023)
by: Cho, Cheol Jun, et al.
Published: (2023)
Analyzing the Impact of Accent on English Speech: Acoustic and Articulatory Perspectives
by: Premananth, Gowtham, et al.
Published: (2025)
by: Premananth, Gowtham, et al.
Published: (2025)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
by: Koriyama, Tomoki
Published: (2025)
by: Koriyama, Tomoki
Published: (2025)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
by: Fu, Ruibo, et al.
Published: (2024)
by: Fu, Ruibo, et al.
Published: (2024)
Accent Conversion with Articulatory Representations
by: Siriwardena, Yashish M., et al.
Published: (2024)
by: Siriwardena, Yashish M., et al.
Published: (2024)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
by: Chen, Zhengyang, et al.
Published: (2024)
by: Chen, Zhengyang, et al.
Published: (2024)
A Phoneme-Scale Assessment of Multichannel Speech Enhancement Algorithms
by: Monir, Nasser-Eddine, et al.
Published: (2024)
by: Monir, Nasser-Eddine, et al.
Published: (2024)
Phoneme-Level Analysis for Person-of-Interest Speech Deepfake Detection
by: Salvi, Davide, et al.
Published: (2025)
by: Salvi, Davide, et al.
Published: (2025)
Beyond Speaker Identity: Text Guided Target Speech Extraction
by: Huo, Mingyue, et al.
Published: (2025)
by: Huo, Mingyue, et al.
Published: (2025)
Enhancing Speaker-Independent Dysarthric Speech Severity Classification with DSSCNet and Cross-Corpus Adaptation
by: Roy, Arnab Kumar, et al.
Published: (2025)
by: Roy, Arnab Kumar, et al.
Published: (2025)
ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
by: Lee, Jihwan, et al.
Published: (2025)
by: Lee, Jihwan, et al.
Published: (2025)
Articulatory Feature Prediction from Surface EMG during Speech Production
by: Lee, Jihwan, et al.
Published: (2025)
by: Lee, Jihwan, et al.
Published: (2025)
Phonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASR
by: Li, Ziwei, et al.
Published: (2026)
by: Li, Ziwei, et al.
Published: (2026)
TASU: Text-Only Alignment for Speech Understanding
by: Peng, Jing, et al.
Published: (2025)
by: Peng, Jing, et al.
Published: (2025)
Evaluating Multichannel Speech Enhancement Algorithms at the Phoneme Scale Across Genders
by: Monir, Nasser-Eddine, et al.
Published: (2025)
by: Monir, Nasser-Eddine, et al.
Published: (2025)
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
by: Singh, Satwinder, et al.
Published: (2025)
by: Singh, Satwinder, et al.
Published: (2025)
Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
by: Dai, Wang, et al.
Published: (2025)
by: Dai, Wang, et al.
Published: (2025)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
by: Zhang, Bowen, et al.
Published: (2025)
by: Zhang, Bowen, et al.
Published: (2025)
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
by: Gállego, Gerard I., et al.
Published: (2025)
by: Gállego, Gerard I., et al.
Published: (2025)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
by: Yamauchi, Kazuki, et al.
Published: (2024)
by: Yamauchi, Kazuki, et al.
Published: (2024)
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
by: Liu, Yisi, et al.
Published: (2025)
by: Liu, Yisi, et al.
Published: (2025)
Inter-Speaker Relative Cues for Two-Stage Text-Guided Target Speech Extraction
by: Dai, Wang, et al.
Published: (2026)
by: Dai, Wang, et al.
Published: (2026)
Similar Items
-
Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis
by: Demir, Kubilay Can, et al.
Published: (2024) -
The Impact of Speech Anonymization on Pathology and Its Limits
by: Arasteh, Soroosh Tayebi, et al.
Published: (2024) -
Towards a Quantitative Analysis of Coarticulation with a Phoneme-to-Articulatory Model
by: Fan, Chaofei, et al.
Published: (2024) -
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
by: Fujita, Kenichi, et al.
Published: (2024) -
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
by: Yang, Dong, et al.
Published: (2025)