Teaching Machines to Speak Using Articulatory Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Anand, Akshay, Guo, Chenxu, Cho, Cheol Jun, Lian, Jiachen, Anumanchipalli, Gopala |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
Towards Hierarchical Spoken Language Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
HuPER: A Human-Inspired Framework for Phonetic Perception
von: Guo, Chenxu, et al.
Veröffentlicht: (2026)
von: Guo, Chenxu, et al.
Veröffentlicht: (2026)
Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP
von: Liu, Yisi, et al.
Veröffentlicht: (2024)
von: Liu, Yisi, et al.
Veröffentlicht: (2024)
Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
von: Zhou, Xuanru, et al.
Veröffentlicht: (2025)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2025)
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
von: Liu, Yisi, et al.
Veröffentlicht: (2025)
von: Liu, Yisi, et al.
Veröffentlicht: (2025)
Scaling Spoken Language Models with Syllabic Speech Tokenization
von: Lee, Nicholas, et al.
Veröffentlicht: (2025)
von: Lee, Nicholas, et al.
Veröffentlicht: (2025)
Sylber 2.0: A Universal Syllable Embedding
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2026)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2026)
Deep Speech Synthesis from Multimodal Articulatory Representations
von: Wu, Peter, et al.
Veröffentlicht: (2024)
von: Wu, Peter, et al.
Veröffentlicht: (2024)
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
Coding Speech through Vocal Tract Kinematics
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
von: Lian, Jiachen, et al.
Veröffentlicht: (2022)
von: Lian, Jiachen, et al.
Veröffentlicht: (2022)
EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
von: Zhou, Xuanru, et al.
Veröffentlicht: (2025)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2025)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
HASS: Hierarchical Simulation of Logopenic Aphasic Speech for Scalable PPA Detection
von: Li, Harrison, et al.
Veröffentlicht: (2026)
von: Li, Harrison, et al.
Veröffentlicht: (2026)
Audio Texture Manipulation by Exemplar-Based Analogy
von: Cheng, Kan Jen, et al.
Veröffentlicht: (2025)
von: Cheng, Kan Jen, et al.
Veröffentlicht: (2025)
Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
Accent Conversion with Articulatory Representations
von: Siriwardena, Yashish M., et al.
Veröffentlicht: (2024)
von: Siriwardena, Yashish M., et al.
Veröffentlicht: (2024)
Acoustic-to-Articulatory Inversion of Clean Speech Using an MRI-Trained Model
von: Azzouz, Sofiane, et al.
Veröffentlicht: (2026)
von: Azzouz, Sofiane, et al.
Veröffentlicht: (2026)
Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection
von: Guo, Chenxu, et al.
Veröffentlicht: (2025)
von: Guo, Chenxu, et al.
Veröffentlicht: (2025)
Acoustic to Articulatory Speech Inversion for Children with Velopharyngeal Insufficiency
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2025)
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2025)
Enhancing Acoustic-to-Articulatory Speech Inversion by Incorporating Nasality
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2025)
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2025)
Reconstruction of the Complete Vocal Tract Contour Through Acoustic to Articulatory Inversion Using Real-Time MRI Data
von: Azzouz, Sofiane, et al.
Veröffentlicht: (2025)
von: Azzouz, Sofiane, et al.
Veröffentlicht: (2025)
Towards a Quantitative Analysis of Coarticulation with a Phoneme-to-Articulatory Model
von: Fan, Chaofei, et al.
Veröffentlicht: (2024)
von: Fan, Chaofei, et al.
Veröffentlicht: (2024)
Towards EMG-to-Speech with a Necklace Form Factor
von: Wu, Peter, et al.
Veröffentlicht: (2024)
von: Wu, Peter, et al.
Veröffentlicht: (2024)
Speech After Gender: A Trans-Feminine Perspective on Next Steps for Speech Science and Technology
von: Netzorg, Robin, et al.
Veröffentlicht: (2024)
von: Netzorg, Robin, et al.
Veröffentlicht: (2024)
Perceptual Ratings Predict Speech Inversion Articulatory Kinematics in Childhood Speech Sound Disorders
von: Benway, Nina R., et al.
Veröffentlicht: (2025)
von: Benway, Nina R., et al.
Veröffentlicht: (2025)
Analyzing the Impact of Accent on English Speech: Acoustic and Articulatory Perspectives
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
Simulating Articulatory Trajectories with Phonological Feature Interpolation
von: Tandazo, Angelo Ortiz, et al.
Veröffentlicht: (2024)
von: Tandazo, Angelo Ortiz, et al.
Veröffentlicht: (2024)
K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function
von: Li, Shuhe, et al.
Veröffentlicht: (2025)
von: Li, Shuhe, et al.
Veröffentlicht: (2025)
SSDM: Scalable Speech Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection
von: Zhang, Jinming, et al.
Veröffentlicht: (2025)
von: Zhang, Jinming, et al.
Veröffentlicht: (2025)
Quantifying Articulatory Coordination as a Biomarker for Schizophrenia
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
von: Premananth, Gowtham, et al.
Veröffentlicht: (2025)
Articulatory modeling of the S-shaped F2 trajectories observed in Öhman's spectrographic analysis of VCV syllables
von: Berthommier, Frédéric
Veröffentlicht: (2025)
von: Berthommier, Frédéric
Veröffentlicht: (2025)
Ähnliche Einträge
-
Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023) -
Towards Hierarchical Spoken Language Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024) -
HuPER: A Human-Inspired Framework for Phonetic Perception
von: Guo, Chenxu, et al.
Veröffentlicht: (2026) -
Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP
von: Liu, Yisi, et al.
Veröffentlicht: (2024) -
Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
von: Zhou, Xuanru, et al.
Veröffentlicht: (2025)