Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ravenscroft, William, Close, George, Goetze, Stefan, Hain, Thomas, Soleymanpour, Mohammad, Chowdhury, Anurag, Fuhs, Mark C. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hallucination in Perceptual Metric-Driven Speech Enhancement Networks
von: Close, George, et al.
Veröffentlicht: (2024)
von: Close, George, et al.
Veröffentlicht: (2024)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
von: Sutherland, Robert, et al.
Veröffentlicht: (2024)
von: Sutherland, Robert, et al.
Veröffentlicht: (2024)
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
von: Close, George, et al.
Veröffentlicht: (2025)
von: Close, George, et al.
Veröffentlicht: (2025)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models
von: Mogridge, Rhiannon, et al.
Veröffentlicht: (2024)
von: Mogridge, Rhiannon, et al.
Veröffentlicht: (2024)
Investigating Confidence Estimation Measures for Speaker Diarization
von: Chowdhury, Anurag, et al.
Veröffentlicht: (2024)
von: Chowdhury, Anurag, et al.
Veröffentlicht: (2024)
Personalized Fine-Tuning with Controllable Synthetic Speech from LLM-Generated Transcripts for Dysarthric Speech Recognition
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
von: Wagner, Dominik, et al.
Veröffentlicht: (2025)
Single-Microphone Speaker Separation and Voice Activity Detection in Noisy and Reverberant Environments
von: Opochinsky, Renana, et al.
Veröffentlicht: (2024)
von: Opochinsky, Renana, et al.
Veröffentlicht: (2024)
Enhanced Reverberation as Supervision for Unsupervised Speech Separation
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
von: von Neumann, Thilo, et al.
Veröffentlicht: (2023)
von: von Neumann, Thilo, et al.
Veröffentlicht: (2023)
Serialized Speech Information Guidance with Overlapped Encoding Separation for Multi-Speaker Automatic Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2024)
von: Shi, Hao, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
von: Park, Chanho, et al.
Veröffentlicht: (2024)
von: Park, Chanho, et al.
Veröffentlicht: (2024)
Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
von: Poncelet, Jakob, et al.
Veröffentlicht: (2025)
von: Poncelet, Jakob, et al.
Veröffentlicht: (2025)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
von: Yang, Yufeng, et al.
Veröffentlicht: (2025)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
von: Li, Guinan, et al.
Veröffentlicht: (2024)
von: Li, Guinan, et al.
Veröffentlicht: (2024)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification
von: Ravenscroft, William, et al.
Veröffentlicht: (2025)
von: Ravenscroft, William, et al.
Veröffentlicht: (2025)
Rehearsal-Free Online Continual Learning for Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2023)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2023)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
ConSep: a Noise- and Reverberation-Robust Speech Separation Framework by Magnitude Conditioning
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
Gradient Norm-based Fine-Tuning for Backdoor Defense in Automatic Speech Recognition
von: Zhou, Nanjun, et al.
Veröffentlicht: (2025)
von: Zhou, Nanjun, et al.
Veröffentlicht: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
von: Tian, Jingguang, et al.
Veröffentlicht: (2024)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
von: Wang, Yuzhu, et al.
Veröffentlicht: (2025)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
von: Huang, Ziling, et al.
Veröffentlicht: (2025)
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghane, Mohsen, et al.
Veröffentlicht: (2025)
Efficient Adapter Tuning of Pre-trained Speech Models for Automatic Speaker Verification
von: Sang, Mufan, et al.
Veröffentlicht: (2024)
von: Sang, Mufan, et al.
Veröffentlicht: (2024)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
von: HU, Shujie, et al.
Veröffentlicht: (2025)
von: HU, Shujie, et al.
Veröffentlicht: (2025)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
von: Alsayegh, Ali, et al.
Veröffentlicht: (2025)
Persian Speech Emotion Recognition by Fine-Tuning Transformers
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
von: Shayaninasab, Minoo, et al.
Veröffentlicht: (2024)
Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
von: Chen, Shuangyuan, et al.
Veröffentlicht: (2025)
von: Chen, Shuangyuan, et al.
Veröffentlicht: (2025)
Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
von: Jing, Ruihao, et al.
Veröffentlicht: (2025)
von: Jing, Ruihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hallucination in Perceptual Metric-Driven Speech Enhancement Networks
von: Close, George, et al.
Veröffentlicht: (2024) -
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
von: Sutherland, Robert, et al.
Veröffentlicht: (2024) -
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
von: Close, George, et al.
Veröffentlicht: (2025) -
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024) -
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models
von: Mogridge, Rhiannon, et al.
Veröffentlicht: (2024)