Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cho, Cheol Jun, Mohamed, Abdelrahman, Black, Alan W, Anumanchipalli, Gopala K. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Spoken Language Models with Syllabic Speech Tokenization
von: Lee, Nicholas, et al.
Veröffentlicht: (2025)
von: Lee, Nicholas, et al.
Veröffentlicht: (2025)
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
Sylber 2.0: A Universal Syllable Embedding
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2026)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2026)
Coding Speech through Vocal Tract Kinematics
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
Teaching Machines to Speak Using Articulatory Control
von: Anand, Akshay, et al.
Veröffentlicht: (2025)
von: Anand, Akshay, et al.
Veröffentlicht: (2025)
Towards Hierarchical Spoken Language Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
Deep Speech Synthesis from Multimodal Articulatory Representations
von: Wu, Peter, et al.
Veröffentlicht: (2024)
von: Wu, Peter, et al.
Veröffentlicht: (2024)
Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP
von: Liu, Yisi, et al.
Veröffentlicht: (2024)
von: Liu, Yisi, et al.
Veröffentlicht: (2024)
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
von: Liu, Yisi, et al.
Veröffentlicht: (2025)
von: Liu, Yisi, et al.
Veröffentlicht: (2025)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
von: Zhou, Xuanru, et al.
Veröffentlicht: (2025)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2025)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
Simulating Articulatory Trajectories with Phonological Feature Interpolation
von: Tandazo, Angelo Ortiz, et al.
Veröffentlicht: (2024)
von: Tandazo, Angelo Ortiz, et al.
Veröffentlicht: (2024)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Tracking Articulatory Dynamics in Speech with a Fixed-Weight BiLSTM-CNN Architecture
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
Acoustic to Articulatory Inversion of Speech; Data Driven Approaches, Challenges, Applications, and Future Scope
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
von: Pillai, Leena G, et al.
Veröffentlicht: (2025)
ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
Interface Design for Self-Supervised Speech Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
Perceptual Ratings Predict Speech Inversion Articulatory Kinematics in Childhood Speech Sound Disorders
von: Benway, Nina R., et al.
Veröffentlicht: (2025)
von: Benway, Nina R., et al.
Veröffentlicht: (2025)
SSDM: Scalable Speech Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
von: Blandón, María Andrea Cruz, et al.
Veröffentlicht: (2025)
von: Blandón, María Andrea Cruz, et al.
Veröffentlicht: (2025)
Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
von: van Rensburg, Kyle Janse, et al.
Veröffentlicht: (2026)
von: van Rensburg, Kyle Janse, et al.
Veröffentlicht: (2026)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
ReHear: Iterative Pseudo-Label Refinement for Semi-Supervised Speech Recognition via Audio Large Language Models
von: Liu, Zefang, et al.
Veröffentlicht: (2026)
von: Liu, Zefang, et al.
Veröffentlicht: (2026)
Towards EMG-to-Speech with a Necklace Form Factor
von: Wu, Peter, et al.
Veröffentlicht: (2024)
von: Wu, Peter, et al.
Veröffentlicht: (2024)
Scaling Self-Supervised Speech Models Uncovers Deep Linguistic Relationships: Evidence from the Pacific Cluster
von: Kim, Minu, et al.
Veröffentlicht: (2026)
von: Kim, Minu, et al.
Veröffentlicht: (2026)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
von: Jang, Kangwook, et al.
Veröffentlicht: (2023)
von: Jang, Kangwook, et al.
Veröffentlicht: (2023)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
von: Lian, Jiachen, et al.
Veröffentlicht: (2022)
von: Lian, Jiachen, et al.
Veröffentlicht: (2022)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2023)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2023)
MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
HuPER: A Human-Inspired Framework for Phonetic Perception
von: Guo, Chenxu, et al.
Veröffentlicht: (2026)
von: Guo, Chenxu, et al.
Veröffentlicht: (2026)
EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Scaling Spoken Language Models with Syllabic Speech Tokenization
von: Lee, Nicholas, et al.
Veröffentlicht: (2025) -
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023) -
Sylber 2.0: A Universal Syllable Embedding
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2026) -
Coding Speech through Vocal Tract Kinematics
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024) -
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)