Inclusive ASR for Disfluent Speech: Cascaded Large-Scale Self-Supervised Learning with Targeted Fine-Tuning and Data Augmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | Mujtaba, Dena, Mahapatra, Nihar R., Arney, Megan, Yaruss, J. Scott, Herring, Caryn, Bin, Jia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Lost in Transcription: Identifying and Quantifying the Accuracy Biases of Automatic Speech Recognition Systems Against Disfluent Speech
di: Mujtaba, Dena, et al.
Pubblicazione: (2024)
di: Mujtaba, Dena, et al.
Pubblicazione: (2024)
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
di: Mujtaba, Dena, et al.
Pubblicazione: (2025)
di: Mujtaba, Dena, et al.
Pubblicazione: (2025)
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
di: Guo, Xin, et al.
Pubblicazione: (2026)
di: Guo, Xin, et al.
Pubblicazione: (2026)
Systematic FAIRness Assessment of Open Voice Biomarker Datasets for Mental Health and Neurodegenerative Diseases
di: Mahapatra, Ishaan, et al.
Pubblicazione: (2025)
di: Mahapatra, Ishaan, et al.
Pubblicazione: (2025)
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
di: Girish, et al.
Pubblicazione: (2026)
di: Girish, et al.
Pubblicazione: (2026)
Bridging Attribution and Open-Set Detection using Graph-Augmented Instance Learning in Synthetic Speech
di: Akhtar, Mohd Mujtaba, et al.
Pubblicazione: (2026)
di: Akhtar, Mohd Mujtaba, et al.
Pubblicazione: (2026)
An approach to measuring the performance of Automatic Speech Recognition (ASR) models in the context of Large Language Model (LLM) powered applications
di: Pulikodan, Sujith, et al.
Pubblicazione: (2025)
di: Pulikodan, Sujith, et al.
Pubblicazione: (2025)
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
di: Kumar, Shashi, et al.
Pubblicazione: (2024)
di: Kumar, Shashi, et al.
Pubblicazione: (2024)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
di: Wang, Huimeng, et al.
Pubblicazione: (2024)
di: Wang, Huimeng, et al.
Pubblicazione: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
di: Shi, Mohan, et al.
Pubblicazione: (2025)
di: Shi, Mohan, et al.
Pubblicazione: (2025)
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
di: Joshi, Sakshi, et al.
Pubblicazione: (2025)
di: Joshi, Sakshi, et al.
Pubblicazione: (2025)
Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
di: Ma, Hao, et al.
Pubblicazione: (2025)
di: Ma, Hao, et al.
Pubblicazione: (2025)
Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR
di: Wang, Zhong-Qiu, et al.
Pubblicazione: (2025)
di: Wang, Zhong-Qiu, et al.
Pubblicazione: (2025)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
ASR for Affective Speech: Investigating Impact of Emotion and Speech Generative Strategy
di: Wu, Ya-Tse, et al.
Pubblicazione: (2026)
di: Wu, Ya-Tse, et al.
Pubblicazione: (2026)
Towards Supervised Performance on Speaker Verification with Self-Supervised Learning by Leveraging Large-Scale ASR Models
di: Miara, Victor, et al.
Pubblicazione: (2024)
di: Miara, Victor, et al.
Pubblicazione: (2024)
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
di: Wang, Weiqing, et al.
Pubblicazione: (2025)
di: Wang, Weiqing, et al.
Pubblicazione: (2025)
BrainWhisperer: Leveraging Large-Scale ASR Models for Neural Speech Decoding
di: Boccato, Tommaso, et al.
Pubblicazione: (2026)
di: Boccato, Tommaso, et al.
Pubblicazione: (2026)
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
di: Menon, Aditya Srinivas, et al.
Pubblicazione: (2026)
di: Menon, Aditya Srinivas, et al.
Pubblicazione: (2026)
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
di: Chung, Raymond
Pubblicazione: (2026)
di: Chung, Raymond
Pubblicazione: (2026)
Rethinking Mamba in Speech Processing by Self-Supervised Models
di: Zhang, Xiangyu, et al.
Pubblicazione: (2024)
di: Zhang, Xiangyu, et al.
Pubblicazione: (2024)
A Unified Denoising and Adaptation Framework for Self-Supervised Bengali Dialectal ASR
di: Biswas, Swadhin, et al.
Pubblicazione: (2025)
di: Biswas, Swadhin, et al.
Pubblicazione: (2025)
Augmenting Open-Vocabulary Dysarthric Speech Assessment with Human Perceptual Supervision
di: Jia, Kaimeng, et al.
Pubblicazione: (2025)
di: Jia, Kaimeng, et al.
Pubblicazione: (2025)
Target Speaker ASR with Whisper
di: Polok, Alexander, et al.
Pubblicazione: (2024)
di: Polok, Alexander, et al.
Pubblicazione: (2024)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
di: Li, Chin-Jou, et al.
Pubblicazione: (2025)
di: Li, Chin-Jou, et al.
Pubblicazione: (2025)
Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis
di: Sohn, Samuel S., et al.
Pubblicazione: (2025)
di: Sohn, Samuel S., et al.
Pubblicazione: (2025)
Metadata-Enhanced Speech Emotion Recognition: Augmented Residual Integration and Co-Attention in Two-Stage Fine-Tuning
di: Wan, Zixiang, et al.
Pubblicazione: (2024)
di: Wan, Zixiang, et al.
Pubblicazione: (2024)
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
di: Yang, Chih-Kai, et al.
Pubblicazione: (2023)
di: Yang, Chih-Kai, et al.
Pubblicazione: (2023)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
Towards Attribution of Generators and Emotional Manipulation in Cross-Lingual Synthetic Speech using Geometric Learning
di: Girish, et al.
Pubblicazione: (2025)
di: Girish, et al.
Pubblicazione: (2025)
ARTT: Augmented Reverberant-Target Training for Unsupervised Monaural Speech Dereverberation
di: Song, Siqi, et al.
Pubblicazione: (2026)
di: Song, Siqi, et al.
Pubblicazione: (2026)
Curved Worlds, Clear Boundaries: Generalizing Speech Deepfake Detection using Hyperbolic and Spherical Geometry Spaces
di: Sheth, Farhan, et al.
Pubblicazione: (2025)
di: Sheth, Farhan, et al.
Pubblicazione: (2025)
Towards a Single ASR Model That Generalizes to Disordered Speech
di: Tobin, Jimmy, et al.
Pubblicazione: (2024)
di: Tobin, Jimmy, et al.
Pubblicazione: (2024)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
di: Chiu, Aemon Yat Fei, et al.
Pubblicazione: (2025)
di: Chiu, Aemon Yat Fei, et al.
Pubblicazione: (2025)
Improving Code Switching with Supervised Fine Tuning and GELU Adapters
di: Pham, Linh
Pubblicazione: (2025)
di: Pham, Linh
Pubblicazione: (2025)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
di: Li, Jialu, et al.
Pubblicazione: (2024)
di: Li, Jialu, et al.
Pubblicazione: (2024)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
di: Cohen, Eyal, et al.
Pubblicazione: (2025)
di: Cohen, Eyal, et al.
Pubblicazione: (2025)
Exploiting Noise Inseparability for Weakly-Supervised Discriminative Speech Denoising Using Noisy Targets
di: Maciejewski, Matthew, et al.
Pubblicazione: (2026)
di: Maciejewski, Matthew, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Lost in Transcription: Identifying and Quantifying the Accuracy Biases of Automatic Speech Recognition Systems Against Disfluent Speech
di: Mujtaba, Dena, et al.
Pubblicazione: (2024) -
Fine-Tuning ASR for Stuttered Speech: Personalized vs. Generalized Approaches
di: Mujtaba, Dena, et al.
Pubblicazione: (2025) -
Adaptive Federated Fine-Tuning of Self-Supervised Speech Representations
di: Guo, Xin, et al.
Pubblicazione: (2026) -
Systematic FAIRness Assessment of Open Voice Biomarker Datasets for Mental Health and Neurodegenerative Diseases
di: Mahapatra, Ishaan, et al.
Pubblicazione: (2025) -
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
di: Girish, et al.
Pubblicazione: (2026)