Sylber 2.0: A Universal Syllable Embedding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cho, Cheol Jun, Lee, Nicholas, Black, Alan W, Anumanchipalli, Gopala K. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
Scaling Spoken Language Models with Syllabic Speech Tokenization
von: Lee, Nicholas, et al.
Veröffentlicht: (2025)
von: Lee, Nicholas, et al.
Veröffentlicht: (2025)
Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
Towards Hierarchical Spoken Language Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
Coding Speech through Vocal Tract Kinematics
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024)
Teaching Machines to Speak Using Articulatory Control
von: Anand, Akshay, et al.
Veröffentlicht: (2025)
von: Anand, Akshay, et al.
Veröffentlicht: (2025)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings
von: Mondal, Anindita, et al.
Veröffentlicht: (2024)
von: Mondal, Anindita, et al.
Veröffentlicht: (2024)
SyllableLM: Learning Coarse Semantic Units for Speech Language Models
von: Baade, Alan, et al.
Veröffentlicht: (2024)
von: Baade, Alan, et al.
Veröffentlicht: (2024)
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
von: Visser, Nicol, et al.
Veröffentlicht: (2026)
von: Visser, Nicol, et al.
Veröffentlicht: (2026)
HuPER: A Human-Inspired Framework for Phonetic Perception
von: Guo, Chenxu, et al.
Veröffentlicht: (2026)
von: Guo, Chenxu, et al.
Veröffentlicht: (2026)
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
von: Komatsu, Ryota, et al.
Veröffentlicht: (2024)
von: Komatsu, Ryota, et al.
Veröffentlicht: (2024)
EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
von: Lian, Jiachen, et al.
Veröffentlicht: (2022)
von: Lian, Jiachen, et al.
Veröffentlicht: (2022)
Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP
von: Liu, Yisi, et al.
Veröffentlicht: (2024)
von: Liu, Yisi, et al.
Veröffentlicht: (2024)
SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
von: Chen, William, et al.
Veröffentlicht: (2025)
von: Chen, William, et al.
Veröffentlicht: (2025)
Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
von: Zhou, Xuanru, et al.
Veröffentlicht: (2025)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2025)
Deep Speech Synthesis from Multimodal Articulatory Representations
von: Wu, Peter, et al.
Veröffentlicht: (2024)
von: Wu, Peter, et al.
Veröffentlicht: (2024)
TalTech Systems for the Interspeech 2025 ML-SUPERB 2.0 Challenge
von: Alumäe, Tanel, et al.
Veröffentlicht: (2025)
von: Alumäe, Tanel, et al.
Veröffentlicht: (2025)
Multimodal Segmentation for Vocal Tract Modeling
von: Jain, Rishi, et al.
Veröffentlicht: (2024)
von: Jain, Rishi, et al.
Veröffentlicht: (2024)
Towards EMG-to-Speech with a Necklace Form Factor
von: Wu, Peter, et al.
Veröffentlicht: (2024)
von: Wu, Peter, et al.
Veröffentlicht: (2024)
Automatic classification of stop realisation with wav2vec2.0
von: Tanner, James, et al.
Veröffentlicht: (2025)
von: Tanner, James, et al.
Veröffentlicht: (2025)
Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling
von: Huang, Shao-Syuan, et al.
Veröffentlicht: (2024)
von: Huang, Shao-Syuan, et al.
Veröffentlicht: (2024)
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
von: Bayerl, Sebastian P., et al.
Veröffentlicht: (2022)
von: Bayerl, Sebastian P., et al.
Veröffentlicht: (2022)
Audio Texture Manipulation by Exemplar-Based Analogy
von: Cheng, Kan Jen, et al.
Veröffentlicht: (2025)
von: Cheng, Kan Jen, et al.
Veröffentlicht: (2025)
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
SSDM: Scalable Speech Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation
von: Tsai, Yun-Shao, et al.
Veröffentlicht: (2026)
von: Tsai, Yun-Shao, et al.
Veröffentlicht: (2026)
Enhancing and Exploring Mild Cognitive Impairment Detection with W2V-BERT-2.0
von: Wang, Yueguan, et al.
Veröffentlicht: (2025)
von: Wang, Yueguan, et al.
Veröffentlicht: (2025)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function
von: Li, Shuhe, et al.
Veröffentlicht: (2025)
von: Li, Shuhe, et al.
Veröffentlicht: (2025)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
von: Huo, Robin, et al.
Veröffentlicht: (2025)
von: Huo, Robin, et al.
Veröffentlicht: (2025)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2024) -
Scaling Spoken Language Models with Syllabic Speech Tokenization
von: Lee, Nicholas, et al.
Veröffentlicht: (2025) -
Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023) -
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023) -
Towards Hierarchical Spoken Language Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)