Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bayerl, Sebastian P., Wagner, Dominik, Nöth, Elmar, Riedhammer, Korbinian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large Language Models for Dysfluency Detection in Stuttered Speech
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Automatic classification of stop realisation with wav2vec2.0
von: Tanner, James, et al.
Veröffentlicht: (2025)
von: Tanner, James, et al.
Veröffentlicht: (2025)
Multilingual Stutter Event Detection for English, German, and Mandarin Speech
von: Haas, Felix, et al.
Veröffentlicht: (2026)
von: Haas, Felix, et al.
Veröffentlicht: (2026)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
von: Gulzar, Kashaf, et al.
Veröffentlicht: (2025)
von: Gulzar, Kashaf, et al.
Veröffentlicht: (2025)
Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
von: Huo, Robin, et al.
Veröffentlicht: (2025)
von: Huo, Robin, et al.
Veröffentlicht: (2025)
Infusing Acoustic Pause Context into Text-Based Dementia Assessment
von: Braun, Franziska, et al.
Veröffentlicht: (2024)
von: Braun, Franziska, et al.
Veröffentlicht: (2024)
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
StutterCut: Uncertainty-Guided Normalised Cut for Dysfluency Segmentation
von: Ghosh, Suhita, et al.
Veröffentlicht: (2025)
von: Ghosh, Suhita, et al.
Veröffentlicht: (2025)
Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation
von: Zhu, Qiushi, et al.
Veröffentlicht: (2024)
von: Zhu, Qiushi, et al.
Veröffentlicht: (2024)
Experimental Study: Enhancing Voice Spoofing Detection Models with wav2vec 2.0
von: Kang, Taein, et al.
Veröffentlicht: (2024)
von: Kang, Taein, et al.
Veröffentlicht: (2024)
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
von: Li, Feng, et al.
Veröffentlicht: (2024)
von: Li, Feng, et al.
Veröffentlicht: (2024)
Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
von: Simic, Christopher, et al.
Veröffentlicht: (2025)
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
MMSD-Net: Towards Multi-modal Stuttering Detection
von: Nie, Liangyu, et al.
Veröffentlicht: (2024)
von: Nie, Liangyu, et al.
Veröffentlicht: (2024)
SSDM: Scalable Speech Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
The PARLO Dementia Corpus: A German Multi-Center Resource for Alzheimer's Disease
von: Braun, Franziska, et al.
Veröffentlicht: (2026)
von: Braun, Franziska, et al.
Veröffentlicht: (2026)
Enhancing and Exploring Mild Cognitive Impairment Detection with W2V-BERT-2.0
von: Wang, Yueguan, et al.
Veröffentlicht: (2025)
von: Wang, Yueguan, et al.
Veröffentlicht: (2025)
Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0
von: Engert, Natalie, et al.
Veröffentlicht: (2026)
von: Engert, Natalie, et al.
Veröffentlicht: (2026)
FGCL: Fine-grained Contrastive Learning For Mandarin Stuttering Event Detection
von: Jiang, Han, et al.
Veröffentlicht: (2024)
von: Jiang, Han, et al.
Veröffentlicht: (2024)
YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
Bias and Fairness in Self-Supervised Acoustic Representations for Cognitive Impairment Detection
von: Gulzar, Kashaf, et al.
Veröffentlicht: (2026)
von: Gulzar, Kashaf, et al.
Veröffentlicht: (2026)
voc2vec: A Foundation Model for Non-Verbal Vocalization
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
von: Xue, Hongfei, et al.
Veröffentlicht: (2024)
Self-supervised Speech Models for Word-Level Stuttered Speech Detection
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection
von: Zhang, Jinming, et al.
Veröffentlicht: (2025)
von: Zhang, Jinming, et al.
Veröffentlicht: (2025)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
von: Huang, Shangkun, et al.
Veröffentlicht: (2025)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
wav2pos: Sound Source Localization using Masked Autoencoders
von: Berg, Axel, et al.
Veröffentlicht: (2024)
von: Berg, Axel, et al.
Veröffentlicht: (2024)
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
von: Mujtaba, Dena, et al.
Veröffentlicht: (2025)
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Large Language Models for Dysfluency Detection in Stuttered Speech
von: Wagner, Dominik, et al.
Veröffentlicht: (2024) -
Automatic classification of stop realisation with wav2vec2.0
von: Tanner, James, et al.
Veröffentlicht: (2025) -
Multilingual Stutter Event Detection for English, German, and Mandarin Speech
von: Haas, Felix, et al.
Veröffentlicht: (2026) -
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
von: Wagner, Dominik, et al.
Veröffentlicht: (2025) -
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)