LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Zongli, Lian, Jiachen, Gupta, Akshaj, Zhou, Xuanru, Li, Haodong, Patel, Krish, Park, Hwi Joo, Zhou, Dingkun, Guo, Chenxu, Li, Shuhe, Wang, Sam, Zhou, Iris, Cho, Cheol Jun, Ezzes, Zoe, Vonk, Jet M. J., Morin, Brittany T., Bogley, Rian, Wauters, Lisa, Miller, Zachary A., Gorno-Tempini, Maria Luisa, Anumanchipalli, Gopala |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
by: Zhou, Xuanru, et al.
Published: (2025)
by: Zhou, Xuanru, et al.
Published: (2025)
Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection
by: Guo, Chenxu, et al.
Published: (2025)
by: Guo, Chenxu, et al.
Published: (2025)
Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection
by: Zhang, Jinming, et al.
Published: (2025)
by: Zhang, Jinming, et al.
Published: (2025)
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis
by: Ye, Zongli, et al.
Published: (2025)
by: Ye, Zongli, et al.
Published: (2025)
SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies
by: Lian, Jiachen, et al.
Published: (2024)
by: Lian, Jiachen, et al.
Published: (2024)
SSDM: Scalable Speech Dysfluency Modeling
by: Lian, Jiachen, et al.
Published: (2024)
by: Lian, Jiachen, et al.
Published: (2024)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
by: Zhou, Xuanru, et al.
Published: (2024)
by: Zhou, Xuanru, et al.
Published: (2024)
Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
by: Zhou, Xuanru, et al.
Published: (2024)
by: Zhou, Xuanru, et al.
Published: (2024)
YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
by: Zhou, Xuanru, et al.
Published: (2024)
by: Zhou, Xuanru, et al.
Published: (2024)
HASS: Hierarchical Simulation of Logopenic Aphasic Speech for Scalable PPA Detection
by: Li, Harrison, et al.
Published: (2026)
by: Li, Harrison, et al.
Published: (2026)
HuPER: A Human-Inspired Framework for Phonetic Perception
by: Guo, Chenxu, et al.
Published: (2026)
by: Guo, Chenxu, et al.
Published: (2026)
Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
by: Zhou, Xuanru, et al.
Published: (2025)
by: Zhou, Xuanru, et al.
Published: (2025)
K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function
by: Li, Shuhe, et al.
Published: (2025)
by: Li, Shuhe, et al.
Published: (2025)
Cognitive reserve and longitudinal changes in brain and cognition in semantic variant primary progressive aphasia
by: Lauren A. Grebe, et al.
Published: (2026)
by: Lauren A. Grebe, et al.
Published: (2026)
Teaching Machines to Speak Using Articulatory Control
by: Anand, Akshay, et al.
Published: (2025)
by: Anand, Akshay, et al.
Published: (2025)
Backward Construct Mapping in Practice: Using Rasch Measurement to Reveal the Internal Structure of Inflectional Morphological Processing in Students with Dyslexia
by: Robin Irey, et al.
Published: (2026)
by: Robin Irey, et al.
Published: (2026)
TART: A Comprehensive Tool for Technique-Aware Audio-to-Tab Guitar Transcription
by: Gupta, Akshaj, et al.
Published: (2025)
by: Gupta, Akshaj, et al.
Published: (2025)
Automated item‐level measures of verbal fluency in semantic and logopenic primary progressive aphasia
by: Jet M. J. Vonk, et al.
Published: (2026)
by: Jet M. J. Vonk, et al.
Published: (2026)
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
by: Zhou, Dingkun, et al.
Published: (2025)
by: Zhou, Dingkun, et al.
Published: (2025)
Time to align sensitive cognitive assessment with protein biomarkers in Alzheimer's disease
by: Jet M. J. Vonk
Published: (2025)
by: Jet M. J. Vonk
Published: (2025)
Asymmetric Hierarchical Anchoring for Audio-Visual Joint Representation: Resolving Information Allocation Ambiguity for Robust Cross-Modal Generalization
by: Wu, Bixing, et al.
Published: (2026)
by: Wu, Bixing, et al.
Published: (2026)
Towards Hierarchical Spoken Language Dysfluency Modeling
by: Lian, Jiachen, et al.
Published: (2024)
by: Lian, Jiachen, et al.
Published: (2024)
Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
by: Cho, Cheol Jun, et al.
Published: (2023)
by: Cho, Cheol Jun, et al.
Published: (2023)
Sylber 2.0: A Universal Syllable Embedding
by: Cho, Cheol Jun, et al.
Published: (2026)
by: Cho, Cheol Jun, et al.
Published: (2026)
Scaling Spoken Language Models with Syllabic Speech Tokenization
by: Lee, Nicholas, et al.
Published: (2025)
by: Lee, Nicholas, et al.
Published: (2025)
A Unified Framework for Model Editing
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
by: Yoon, Junsang, et al.
Published: (2024)
by: Yoon, Junsang, et al.
Published: (2024)
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
StyleStream: Real-Time Zero-Shot Voice Style Conversion
by: Liu, Yisi, et al.
Published: (2026)
by: Liu, Yisi, et al.
Published: (2026)
Self-Assessment Tests are Unreliable Measures of LLM Personality
by: Gupta, Akshat, et al.
Published: (2023)
by: Gupta, Akshat, et al.
Published: (2023)
Coding Speech through Vocal Tract Kinematics
by: Cho, Cheol Jun, et al.
Published: (2024)
by: Cho, Cheol Jun, et al.
Published: (2024)
Combining script training and lexical retrieval treatment with remotely supervised transcranial direct current stimulation in primary progressive aphasia
by: Lisa D. Wauters, et al.
Published: (2025)
by: Lisa D. Wauters, et al.
Published: (2025)
Audio Texture Manipulation by Exemplar-Based Analogy
by: Cheng, Kan Jen, et al.
Published: (2025)
by: Cheng, Kan Jen, et al.
Published: (2025)
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
by: Cho, Cheol Jun, et al.
Published: (2023)
by: Cho, Cheol Jun, et al.
Published: (2023)
Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP
by: Liu, Yisi, et al.
Published: (2024)
by: Liu, Yisi, et al.
Published: (2024)
How Do LLMs Use Their Depth?
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Efficient Knowledge Editing via Minimal Precomputation
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
by: Lian, Jiachen, et al.
Published: (2022)
by: Lian, Jiachen, et al.
Published: (2022)
Similar Items
-
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
by: Zhou, Xuanru, et al.
Published: (2025) -
Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection
by: Guo, Chenxu, et al.
Published: (2025) -
Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection
by: Zhang, Jinming, et al.
Published: (2025) -
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis
by: Ye, Zongli, et al.
Published: (2025) -
SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies
by: Lian, Jiachen, et al.
Published: (2024)