Heterogeneous Space Fusion and Dual-Dimension Attention: A New Paradigm for Speech Enhancement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Tao, Wang, Liejun, Yu, Yinfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network based on Self-Supervised Embedding
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2024)
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2024)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders
von: Cao, Yubing, et al.
Veröffentlicht: (2024)
von: Cao, Yubing, et al.
Veröffentlicht: (2024)
ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
von: Zhu, Tao, et al.
Veröffentlicht: (2025)
von: Zhu, Tao, et al.
Veröffentlicht: (2025)
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
von: Cao, Yubing, et al.
Veröffentlicht: (2025)
von: Cao, Yubing, et al.
Veröffentlicht: (2025)
ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations
von: Wang, Kexue, et al.
Veröffentlicht: (2026)
von: Wang, Kexue, et al.
Veröffentlicht: (2026)
Leveraging Label Potential for Enhanced Multimodal Emotion Recognition
von: Shao, Xuechun, et al.
Veröffentlicht: (2025)
von: Shao, Xuechun, et al.
Veröffentlicht: (2025)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
Audio Spatially-Guided Fusion for Audio-Visual Navigation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
Spatial-Aware Conditioned Fusion for Audio-Visual Navigation
von: Wu, Shaohang, et al.
Veröffentlicht: (2026)
von: Wu, Shaohang, et al.
Veröffentlicht: (2026)
Reliability-Aware Geometric Fusion for Robust Audio-Visual Navigation
von: Liu, Teng, et al.
Veröffentlicht: (2026)
von: Liu, Teng, et al.
Veröffentlicht: (2026)
Generalizable Audio-Visual Navigation via Binaural Difference Attention and Action Transition Prediction
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
PCQ: Emotion Recognition in Speech via Progressive Channel Querying
von: Wang, Xincheng, et al.
Veröffentlicht: (2024)
von: Wang, Xincheng, et al.
Veröffentlicht: (2024)
Visual-Informed Speech Enhancement Using Attention-Based Beamforming
von: Liu, Chihyun, et al.
Veröffentlicht: (2026)
von: Liu, Chihyun, et al.
Veröffentlicht: (2026)
A Dual-Branch Parallel Network for Speech Enhancement and Restoration
von: Yang, Da-Hee, et al.
Veröffentlicht: (2024)
von: Yang, Da-Hee, et al.
Veröffentlicht: (2024)
A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis
von: Hu, Guoqiang, et al.
Veröffentlicht: (2024)
von: Hu, Guoqiang, et al.
Veröffentlicht: (2024)
Modality-Invariant Bidirectional Temporal Representation Distillation Network for Missing Multimodal Sentiment Analysis
von: Wang, Xincheng, et al.
Veröffentlicht: (2025)
von: Wang, Xincheng, et al.
Veröffentlicht: (2025)
MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
Real-Time Speech Enhancement via a Hybrid ViT: A Dual-Input Acoustic-Image Feature Fusion
von: Bahmei, Behnaz, et al.
Veröffentlicht: (2025)
von: Bahmei, Behnaz, et al.
Veröffentlicht: (2025)
EffiFusion-GAN: Efficient Fusion Generative Adversarial Network for Speech Enhancement
von: Wen, Bin, et al.
Veröffentlicht: (2025)
von: Wen, Bin, et al.
Veröffentlicht: (2025)
Exploring Resolution-Wise Shared Attention in Hybrid Mamba-U-Nets for Improved Cross-Corpus Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
von: Yaish, Ofir, et al.
Veröffentlicht: (2025)
von: Yaish, Ofir, et al.
Veröffentlicht: (2025)
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
TS-URGENet: A Three-stage Universal Robust and Generalizable Speech Enhancement Network
von: Rong, Xiaobin, et al.
Veröffentlicht: (2025)
von: Rong, Xiaobin, et al.
Veröffentlicht: (2025)
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement
von: Bae, Jae-Sung, et al.
Veröffentlicht: (2025)
von: Bae, Jae-Sung, et al.
Veröffentlicht: (2025)
An Investigation of Incorporating Mamba for Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2024)
von: Chao, Rong, et al.
Veröffentlicht: (2024)
HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
von: Li, Bohan, et al.
Veröffentlicht: (2026)
von: Li, Bohan, et al.
Veröffentlicht: (2026)
Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
von: Wang, Junyu, et al.
Veröffentlicht: (2024)
von: Wang, Junyu, et al.
Veröffentlicht: (2024)
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
von: Chen, Chih-Ning, et al.
Veröffentlicht: (2026)
von: Chen, Chih-Ning, et al.
Veröffentlicht: (2026)
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance
von: Milling, Manuel, et al.
Veröffentlicht: (2024)
von: Milling, Manuel, et al.
Veröffentlicht: (2024)
Joint Learning using Mixture-of-Expert-Based Representation for Speech Enhancement and Robust Emotion Recognition
von: Tzeng, Jing-Tong, et al.
Veröffentlicht: (2025)
von: Tzeng, Jing-Tong, et al.
Veröffentlicht: (2025)
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
A Survey of Deep Learning for Complex Speech Spectrograms
von: Xie, Yuying, et al.
Veröffentlicht: (2025)
von: Xie, Yuying, et al.
Veröffentlicht: (2025)
Assessment of Personality Dimensions Across Situations Using Conversational Speech
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
von: Lu, Shenghui, et al.
Veröffentlicht: (2025)
von: Lu, Shenghui, et al.
Veröffentlicht: (2025)
Unsupervised Speech Enhancement using Data-defined Priors
von: Klement, Dominik, et al.
Veröffentlicht: (2025)
von: Klement, Dominik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
BSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network based on Self-Supervised Embedding
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2024) -
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024) -
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025) -
VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders
von: Cao, Yubing, et al.
Veröffentlicht: (2024) -
ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
von: Zhu, Tao, et al.
Veröffentlicht: (2025)