YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Xuanru, Kashyap, Anshul, Li, Steve, Sharma, Ayati, Morin, Brittany, Baquirin, David, Vonk, Jet, Ezzes, Zoe, Miller, Zachary, Tempini, Maria Luisa Gorno, Lian, Jiachen, Anumanchipalli, Gopala Krishna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
by: Zhou, Xuanru, et al.
Published: (2024)
by: Zhou, Xuanru, et al.
Published: (2024)
SSDM: Scalable Speech Dysfluency Modeling
by: Lian, Jiachen, et al.
Published: (2024)
by: Lian, Jiachen, et al.
Published: (2024)
SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies
by: Lian, Jiachen, et al.
Published: (2024)
by: Lian, Jiachen, et al.
Published: (2024)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
by: Zhou, Xuanru, et al.
Published: (2024)
by: Zhou, Xuanru, et al.
Published: (2024)
Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection
by: Zhang, Jinming, et al.
Published: (2025)
by: Zhang, Jinming, et al.
Published: (2025)
Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection
by: Guo, Chenxu, et al.
Published: (2025)
by: Guo, Chenxu, et al.
Published: (2025)
Towards Hierarchical Spoken Language Dysfluency Modeling
by: Lian, Jiachen, et al.
Published: (2024)
by: Lian, Jiachen, et al.
Published: (2024)
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
by: Zhou, Xuanru, et al.
Published: (2025)
by: Zhou, Xuanru, et al.
Published: (2025)
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis
by: Ye, Zongli, et al.
Published: (2025)
by: Ye, Zongli, et al.
Published: (2025)
HASS: Hierarchical Simulation of Logopenic Aphasic Speech for Scalable PPA Detection
by: Li, Harrison, et al.
Published: (2026)
by: Li, Harrison, et al.
Published: (2026)
LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
by: Ye, Zongli, et al.
Published: (2025)
by: Ye, Zongli, et al.
Published: (2025)
Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
by: Zhou, Xuanru, et al.
Published: (2025)
by: Zhou, Xuanru, et al.
Published: (2025)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
by: Lian, Jiachen, et al.
Published: (2022)
by: Lian, Jiachen, et al.
Published: (2022)
Backward Construct Mapping in Practice: Using Rasch Measurement to Reveal the Internal Structure of Inflectional Morphological Processing in Students with Dyslexia
by: Robin Irey, et al.
Published: (2026)
by: Robin Irey, et al.
Published: (2026)
Large Language Models for Dysfluency Detection in Stuttered Speech
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
Cognitive reserve and longitudinal changes in brain and cognition in semantic variant primary progressive aphasia
by: Lauren A. Grebe, et al.
Published: (2026)
by: Lauren A. Grebe, et al.
Published: (2026)
StutterCut: Uncertainty-Guided Normalised Cut for Dysfluency Segmentation
by: Ghosh, Suhita, et al.
Published: (2025)
by: Ghosh, Suhita, et al.
Published: (2025)
Teaching Machines to Speak Using Articulatory Control
by: Anand, Akshay, et al.
Published: (2025)
by: Anand, Akshay, et al.
Published: (2025)
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
by: Bayerl, Sebastian P., et al.
Published: (2022)
by: Bayerl, Sebastian P., et al.
Published: (2022)
Time to align sensitive cognitive assessment with protein biomarkers in Alzheimer's disease
by: Jet M. J. Vonk
Published: (2025)
by: Jet M. J. Vonk
Published: (2025)
Asymmetric Hierarchical Anchoring for Audio-Visual Joint Representation: Resolving Information Allocation Ambiguity for Robust Cross-Modal Generalization
by: Wu, Bixing, et al.
Published: (2026)
by: Wu, Bixing, et al.
Published: (2026)
TART: A Comprehensive Tool for Technique-Aware Audio-to-Tab Guitar Transcription
by: Gupta, Akshaj, et al.
Published: (2025)
by: Gupta, Akshaj, et al.
Published: (2025)
A Unified Framework for Model Editing
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
by: Yoon, Junsang, et al.
Published: (2024)
by: Yoon, Junsang, et al.
Published: (2024)
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
StyleStream: Real-Time Zero-Shot Voice Style Conversion
by: Liu, Yisi, et al.
Published: (2026)
by: Liu, Yisi, et al.
Published: (2026)
Self-Assessment Tests are Unreliable Measures of LLM Personality
by: Gupta, Akshat, et al.
Published: (2023)
by: Gupta, Akshat, et al.
Published: (2023)
Automated item‐level measures of verbal fluency in semantic and logopenic primary progressive aphasia
by: Jet M. J. Vonk, et al.
Published: (2026)
by: Jet M. J. Vonk, et al.
Published: (2026)
HuPER: A Human-Inspired Framework for Phonetic Perception
by: Guo, Chenxu, et al.
Published: (2026)
by: Guo, Chenxu, et al.
Published: (2026)
K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function
by: Li, Shuhe, et al.
Published: (2025)
by: Li, Shuhe, et al.
Published: (2025)
Audio Texture Manipulation by Exemplar-Based Analogy
by: Cheng, Kan Jen, et al.
Published: (2025)
by: Cheng, Kan Jen, et al.
Published: (2025)
Tong‐man Ong (June 9, 1935–January 5, 2025)
by: Gopala Krishna
Published: (2025)
by: Gopala Krishna
Published: (2025)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
by: Lin, Guan-Ting, et al.
Published: (2025)
by: Lin, Guan-Ting, et al.
Published: (2025)
Response to rethinking Alzheimer's disease susceptibility and heterogeneity. Comment on Miller et al
by: Zachary A. Miller, et al.
Published: (2026)
by: Zachary A. Miller, et al.
Published: (2026)
Speech After Gender: A Trans-Feminine Perspective on Next Steps for Speech Science and Technology
by: Netzorg, Robin, et al.
Published: (2024)
by: Netzorg, Robin, et al.
Published: (2024)
How Do LLMs Use Their Depth?
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Efficient Knowledge Editing via Minimal Precomputation
by: Gupta, Akshat, et al.
Published: (2025)
by: Gupta, Akshat, et al.
Published: (2025)
Remote Naturalistic Speech Analysis as a Tool to Classify Behavioral Variant Frontotemporal Dementia
by: Jet M. J. Vonk, et al.
Published: (2024)
by: Jet M. J. Vonk, et al.
Published: (2024)
Similar Items
-
Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
by: Zhou, Xuanru, et al.
Published: (2024) -
SSDM: Scalable Speech Dysfluency Modeling
by: Lian, Jiachen, et al.
Published: (2024) -
SSDM 2.0: Time-Accurate Speech Rich Transcription with Non-Fluencies
by: Lian, Jiachen, et al.
Published: (2024) -
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
by: Zhou, Xuanru, et al.
Published: (2024) -
Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection
by: Zhang, Jinming, et al.
Published: (2025)