Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cui, Zhongjian, Cui, Chenrui, Wang, Tianrui, He, Mengnan, Shi, Hao, Ge, Meng, Gong, Caixia, Wang, Longbiao, Dang, Jianwu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
von: Wang, Junyu, et al.
Veröffentlicht: (2024)
von: Wang, Junyu, et al.
Veröffentlicht: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
Progressive Residual Extraction based Pre-training for Speech Representation Learning
von: Wang, Tianrui, et al.
Veröffentlicht: (2024)
von: Wang, Tianrui, et al.
Veröffentlicht: (2024)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
von: Shu, Yuchun, et al.
Veröffentlicht: (2024)
von: Shu, Yuchun, et al.
Veröffentlicht: (2024)
ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
von: Nasretdinov, Rauf, et al.
Veröffentlicht: (2025)
von: Nasretdinov, Rauf, et al.
Veröffentlicht: (2025)
Towards Lightweight Adaptation of Speech Enhancement Models in Real-World Environments
von: Cheng, Longbiao, et al.
Veröffentlicht: (2026)
von: Cheng, Longbiao, et al.
Veröffentlicht: (2026)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
von: Chen, Yanan, et al.
Veröffentlicht: (2024)
von: Chen, Yanan, et al.
Veröffentlicht: (2024)
SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
von: Zhao, Jiahui, et al.
Veröffentlicht: (2024)
von: Zhao, Jiahui, et al.
Veröffentlicht: (2024)
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
Diffusion-based Speech Enhancement with Schrödinger Bridge and Symmetric Noise Schedule
von: Wang, Siyi, et al.
Veröffentlicht: (2024)
von: Wang, Siyi, et al.
Veröffentlicht: (2024)
Schrödinger Bridge Consistency Trajectory Models for Speech Enhancement
von: Nishigori, Shuichiro, et al.
Veröffentlicht: (2025)
von: Nishigori, Shuichiro, et al.
Veröffentlicht: (2025)
VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
von: Qiang, Chunyu, et al.
Veröffentlicht: (2024)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2024)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
von: Ochiai, Tsubasa, et al.
Veröffentlicht: (2024)
von: Ochiai, Tsubasa, et al.
Veröffentlicht: (2024)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
von: HU, Shujie, et al.
Veröffentlicht: (2025)
von: HU, Shujie, et al.
Veröffentlicht: (2025)
Dynamic Gated Recurrent Neural Network for Compute-efficient Speech Enhancement
von: Cheng, Longbiao, et al.
Veröffentlicht: (2024)
von: Cheng, Longbiao, et al.
Veröffentlicht: (2024)
Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement
von: Cheng, Longbiao, et al.
Veröffentlicht: (2024)
von: Cheng, Longbiao, et al.
Veröffentlicht: (2024)
Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
von: Han, Seungu, et al.
Veröffentlicht: (2025)
von: Han, Seungu, et al.
Veröffentlicht: (2025)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
von: Yang, Bing, et al.
Veröffentlicht: (2024)
von: Yang, Bing, et al.
Veröffentlicht: (2024)
SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition
von: Halim, Jule Valendo, et al.
Veröffentlicht: (2025)
von: Halim, Jule Valendo, et al.
Veröffentlicht: (2025)
DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
Restorative Speech Enhancement: A Progressive Approach Using SE and Codec Modules
von: Chiang, Hsin-Tien, et al.
Veröffentlicht: (2024)
von: Chiang, Hsin-Tien, et al.
Veröffentlicht: (2024)
Universal Speech Enhancement with Regression and Generative Mamba
von: Chao, Rong, et al.
Veröffentlicht: (2025)
von: Chao, Rong, et al.
Veröffentlicht: (2025)
In-Materia Speech Recognition
von: Zolfagharinejad, Mohamadreza, et al.
Veröffentlicht: (2024)
von: Zolfagharinejad, Mohamadreza, et al.
Veröffentlicht: (2024)
Reducing Geographic Disparities in Automatic Speech Recognition via Elastic Weight Consolidation
von: Trinh, Viet Anh, et al.
Veröffentlicht: (2022)
von: Trinh, Viet Anh, et al.
Veröffentlicht: (2022)
Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
von: Gao, Yuan, et al.
Veröffentlicht: (2025)
ICASSP 2026 URGENT Speech Enhancement Challenge
von: Li, Chenda, et al.
Veröffentlicht: (2026)
von: Li, Chenda, et al.
Veröffentlicht: (2026)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
von: Ma, Ding, et al.
Veröffentlicht: (2026)
von: Ma, Ding, et al.
Veröffentlicht: (2026)
Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
von: de Groot, Dimme, et al.
Veröffentlicht: (2025)
von: de Groot, Dimme, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
von: Wang, Junyu, et al.
Veröffentlicht: (2025) -
Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
von: Wang, Junyu, et al.
Veröffentlicht: (2024) -
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023) -
Progressive Residual Extraction based Pre-training for Speech Representation Learning
von: Wang, Tianrui, et al.
Veröffentlicht: (2024) -
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)