SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Luo, Longjie, Li, Lin, Hong, Qingyang |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
par: Luo, Longjie, et autres
Publié: (2025)
par: Luo, Longjie, et autres
Publié: (2025)
ctPuLSE: Close-Talk, and Pseudo-Label Based Far-Field, Speech Enhancement
par: Wang, Zhong-Qiu
Publié: (2024)
par: Wang, Zhong-Qiu
Publié: (2024)
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
par: Lin, Zhaofeng, et autres
Publié: (2023)
par: Lin, Zhaofeng, et autres
Publié: (2023)
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
par: Nasretdinov, Rauf, et autres
Publié: (2025)
par: Nasretdinov, Rauf, et autres
Publié: (2025)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
par: Lin, Yi-Cheng, et autres
Publié: (2025)
par: Lin, Yi-Cheng, et autres
Publié: (2025)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
par: Chen, Yanan, et autres
Publié: (2024)
par: Chen, Yanan, et autres
Publié: (2024)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
par: Zezario, Ryandhimas E., et autres
Publié: (2026)
par: Zezario, Ryandhimas E., et autres
Publié: (2026)
Adaptive Data Augmentation with NaturalSpeech3 for Far-field Speaker Verification
par: Zhang, Li, et autres
Publié: (2025)
par: Zhang, Li, et autres
Publié: (2025)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
par: Tsunoo, Emiru, et autres
Publié: (2024)
par: Tsunoo, Emiru, et autres
Publié: (2024)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
par: Ochiai, Tsubasa, et autres
Publié: (2024)
par: Ochiai, Tsubasa, et autres
Publié: (2024)
A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
par: Lu, Shenghui, et autres
Publié: (2025)
par: Lu, Shenghui, et autres
Publié: (2025)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
par: Ma, Ding, et autres
Publié: (2026)
par: Ma, Ding, et autres
Publié: (2026)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
par: Wang, Kuan-Chen, et autres
Publié: (2024)
par: Wang, Kuan-Chen, et autres
Publié: (2024)
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
par: Guan, Wenhao, et autres
Publié: (2023)
par: Guan, Wenhao, et autres
Publié: (2023)
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
par: Huang, Hukai, et autres
Publié: (2024)
par: Huang, Hukai, et autres
Publié: (2024)
GMP-TL: Gender-augmented Multi-scale Pseudo-label Enhanced Transfer Learning for Speech Emotion Recognition
par: Pan, Yu, et autres
Publié: (2024)
par: Pan, Yu, et autres
Publié: (2024)
WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
par: Fang, Zihao, et autres
Publié: (2026)
par: Fang, Zihao, et autres
Publié: (2026)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
par: Yang, Da-Hee, et autres
Publié: (2026)
par: Yang, Da-Hee, et autres
Publié: (2026)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
par: Yen, Hao, et autres
Publié: (2024)
par: Yen, Hao, et autres
Publié: (2024)
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
par: Chen, Peijie, et autres
Publié: (2025)
par: Chen, Peijie, et autres
Publié: (2025)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
par: Guan, Wenhao, et autres
Publié: (2023)
par: Guan, Wenhao, et autres
Publié: (2023)
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
par: Cui, Zhongjian, et autres
Publié: (2025)
par: Cui, Zhongjian, et autres
Publié: (2025)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
par: Gong, Xun, et autres
Publié: (2024)
par: Gong, Xun, et autres
Publié: (2024)
Multichannel Long-Term Streaming Neural Speech Enhancement for Static and Moving Speakers
par: Quan, Changsheng, et autres
Publié: (2024)
par: Quan, Changsheng, et autres
Publié: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
par: Lin, Guan-Ting, et autres
Publié: (2024)
par: Lin, Guan-Ting, et autres
Publié: (2024)
Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
par: Wu, Weijie, et autres
Publié: (2025)
par: Wu, Weijie, et autres
Publié: (2025)
PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data
par: Cao, Songjun, et autres
Publié: (2025)
par: Cao, Songjun, et autres
Publié: (2025)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
par: Ren, Wenze, et autres
Publié: (2024)
par: Ren, Wenze, et autres
Publié: (2024)
Probing Self-supervised Learning Models with Target Speech Extraction
par: Peng, Junyi, et autres
Publié: (2024)
par: Peng, Junyi, et autres
Publié: (2024)
Speech Emotion Recognition with ASR Integration
par: Li, Yuanchao
Publié: (2026)
par: Li, Yuanchao
Publié: (2026)
Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model
par: Zezario, Ryandhimas E., et autres
Publié: (2023)
par: Zezario, Ryandhimas E., et autres
Publié: (2023)
In-Materia Speech Recognition
par: Zolfagharinejad, Mohamadreza, et autres
Publié: (2024)
par: Zolfagharinejad, Mohamadreza, et autres
Publié: (2024)
Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
par: Wen, Wen, et autres
Publié: (2024)
par: Wen, Wen, et autres
Publié: (2024)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
par: Zhu, Xinfa, et autres
Publié: (2023)
par: Zhu, Xinfa, et autres
Publié: (2023)
On the Importance of Neural Wiener Filter for Resource Efficient Multichannel Speech Enhancement
par: Hsieh, Tsun-An, et autres
Publié: (2024)
par: Hsieh, Tsun-An, et autres
Publié: (2024)
Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
par: de Groot, Dimme, et autres
Publié: (2025)
par: de Groot, Dimme, et autres
Publié: (2025)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
par: Lin, Zhaofeng, et autres
Publié: (2024)
par: Lin, Zhaofeng, et autres
Publié: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
par: Shen, Siyuan, et autres
Publié: (2024)
par: Shen, Siyuan, et autres
Publié: (2024)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
par: Meghanani, Amit, et autres
Publié: (2026)
par: Meghanani, Amit, et autres
Publié: (2026)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
par: Peng, Junyi, et autres
Publié: (2024)
par: Peng, Junyi, et autres
Publié: (2024)
Documents similaires
-
Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
par: Luo, Longjie, et autres
Publié: (2025) -
ctPuLSE: Close-Talk, and Pseudo-Label Based Far-Field, Speech Enhancement
par: Wang, Zhong-Qiu
Publié: (2024) -
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
par: Lin, Zhaofeng, et autres
Publié: (2023) -
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
par: Nasretdinov, Rauf, et autres
Publié: (2025) -
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
par: Lin, Yi-Cheng, et autres
Publié: (2025)