SuperM2M: Supervised and Mixture-to-Mixture Co-Learning for Speech Enhancement and Noise-Robust ASR
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Wang, Zhong-Qiu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR
von: Wang, Zhong-Qiu, et al.
Veröffentlicht: (2025)
von: Wang, Zhong-Qiu, et al.
Veröffentlicht: (2025)
Mixture to Mixture: Leveraging Close-talk Mixtures as Weak-supervision for Speech Separation
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
USDnet: Unsupervised Speech Dereverberation via Neural Forward Filtering
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
von: Wang, Zhong-Qiu
Veröffentlicht: (2024)
Toward Universal Speech Enhancement for Diverse Input Conditions
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM
von: Gong, Xun, et al.
Veröffentlicht: (2025)
von: Gong, Xun, et al.
Veröffentlicht: (2025)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Binaural Localization Model for Speech in Noise
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
Speech Enhancement based on cascaded two flows
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
SELM: Speech Enhancement Using Discrete Tokens and Language Models
von: Wang, Ziqian, et al.
Veröffentlicht: (2023)
von: Wang, Ziqian, et al.
Veröffentlicht: (2023)
FlowSE: Flow Matching-based Speech Enhancement
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
von: Zhang, Wangyou, et al.
Veröffentlicht: (2025)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2025)
Binaural Speech Enhancement Using Complex Convolutional Recurrent Networks
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
von: Tokala, Vikas, et al.
Veröffentlicht: (2025)
FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
von: Lee, Dongheon, et al.
Veröffentlicht: (2023)
von: Lee, Dongheon, et al.
Veröffentlicht: (2023)
HyBeam: Hybrid Microphone-Beamforming Array-Agnostic Speech Enhancement for Wearables
von: Ilan, Yuval Bar, et al.
Veröffentlicht: (2025)
von: Ilan, Yuval Bar, et al.
Veröffentlicht: (2025)
Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
von: Aihara, Ryo, et al.
Veröffentlicht: (2025)
Unsupervised Face-Masked Speech Enhancement Using Generative Adversarial Networks With Human-in-the-Loop Assessment Metrics
von: Wang, Syu-Siang, et al.
Veröffentlicht: (2024)
von: Wang, Syu-Siang, et al.
Veröffentlicht: (2024)
A Robust Proactive Communication Strategy for Distributed Active Noise Control Systems
von: Ji, Junwei, et al.
Veröffentlicht: (2025)
von: Ji, Junwei, et al.
Veröffentlicht: (2025)
TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
PLDNet: PLD-Guided Lightweight Deep Network Boosted by Efficient Attention for Handheld Dual-Microphone Speech Enhancement
von: Zhou, Nan, et al.
Veröffentlicht: (2024)
von: Zhou, Nan, et al.
Veröffentlicht: (2024)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2022)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
von: Serre, Thomas, et al.
Veröffentlicht: (2026)
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
Robust Detection of Underwater Target Against Non-Uniform Noise With Optical Fiber DAS Array
von: Cang, Siyuan, et al.
Veröffentlicht: (2025)
von: Cang, Siyuan, et al.
Veröffentlicht: (2025)
Self-Boosted Weight-Constrained FxLMS: A Robustness Distributed Active Noise Control Algorithm Without Internode Communication
von: Ji, Junwei, et al.
Veröffentlicht: (2025)
von: Ji, Junwei, et al.
Veröffentlicht: (2025)
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
von: Masuyama, Yoshiki, et al.
Veröffentlicht: (2025)
Transferable Selective Virtual Sensing Active Noise Control Technique Based on Metric Learning
von: Wang, Boxiang, et al.
Veröffentlicht: (2024)
von: Wang, Boxiang, et al.
Veröffentlicht: (2024)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
SpeechMLC: Speech Multi-label Classification
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
von: Yang, Yujie, et al.
Veröffentlicht: (2025)
von: Yang, Yujie, et al.
Veröffentlicht: (2025)
Impact of Microphone Array Mismatches to Learning-based Replay Speech Detection
von: Neri, Michael, et al.
Veröffentlicht: (2025)
von: Neri, Michael, et al.
Veröffentlicht: (2025)
Distributed Multichannel Active Noise Control with Asynchronous Communication
von: Ji, Junwei, et al.
Veröffentlicht: (2026)
von: Ji, Junwei, et al.
Veröffentlicht: (2026)
Implementation of the Feedforward Multichannel Virtual Sensing Active Noise Control (MVANC) by Using MATLAB
von: Wang, Boxiang
Veröffentlicht: (2024)
von: Wang, Boxiang
Veröffentlicht: (2024)
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2025)
von: Kamo, Naoyuki, et al.
Veröffentlicht: (2025)
Optimizing Domain-Adaptive Self-Supervised Learning for Clinical Voice-Based Disease Classification
von: Liu, Weixin, et al.
Veröffentlicht: (2026)
von: Liu, Weixin, et al.
Veröffentlicht: (2026)
Acoustic Non-Stationarity Objective Assessment with Hard Label Criteria for Supervised Learning Models
von: Zucatelli, Guilherme, et al.
Veröffentlicht: (2025)
von: Zucatelli, Guilherme, et al.
Veröffentlicht: (2025)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
von: Yan, Haoyin, et al.
Veröffentlicht: (2024)
von: Yan, Haoyin, et al.
Veröffentlicht: (2024)
The Overview of Segmental Durations Modification Algorithms on Speech Signal Characteristics
von: Jang, Kyeomeun, et al.
Veröffentlicht: (2025)
von: Jang, Kyeomeun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR
von: Wang, Zhong-Qiu, et al.
Veröffentlicht: (2025) -
Mixture to Mixture: Leveraging Close-talk Mixtures as Weak-supervision for Speech Separation
von: Wang, Zhong-Qiu
Veröffentlicht: (2024) -
USDnet: Unsupervised Speech Dereverberation via Neural Forward Filtering
von: Wang, Zhong-Qiu
Veröffentlicht: (2024) -
Toward Universal Speech Enhancement for Diverse Input Conditions
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023) -
Generic Speech Enhancement with Self-Supervised Representation Space Loss
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)