I-DCCRN-VAE: An Improved Deep Representation Learning Framework for Complex VAE-based Single-channel Speech Enhancement
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jiatong, Doclo, Simon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigation of Speech and Noise Latent Representations in Single-channel VAE-based Speech Enhancement
by: Li, Jiatong, et al.
Published: (2025)
by: Li, Jiatong, et al.
Published: (2025)
Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
by: Han, Runduo, et al.
Published: (2024)
by: Han, Runduo, et al.
Published: (2024)
Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks
by: Tokala, Vikas, et al.
Published: (2024)
by: Tokala, Vikas, et al.
Published: (2024)
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
by: Niu, Zhikang, et al.
Published: (2025)
by: Niu, Zhikang, et al.
Published: (2025)
Flexible Multi-Channel Target Speaker Extraction Using Geometry-Conditioned Spatially Selective Non-linear Filters
by: Li, Jiatong, et al.
Published: (2026)
by: Li, Jiatong, et al.
Published: (2026)
Binaural Speech Enhancement Using Complex Convolutional Recurrent Networks
by: Tokala, Vikas, et al.
Published: (2025)
by: Tokala, Vikas, et al.
Published: (2025)
Comparison of Knowledge Distillation Methods for Low-complexity Multi-microphone Speech Enhancement using the FT-JNF Architecture
by: Metzger, Robert, et al.
Published: (2025)
by: Metzger, Robert, et al.
Published: (2025)
Cross-Utterance Conditioned VAE for Speech Generation
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
Deep low-latency joint speech transmission and enhancement over a gaussian channel
by: Bokaei, Mohammad, et al.
Published: (2024)
by: Bokaei, Mohammad, et al.
Published: (2024)
Speech-dependent Data Augmentation for Own Voice Reconstruction with Hearable Microphones in Noisy Environments
by: Ohlenbusch, Mattes, et al.
Published: (2024)
by: Ohlenbusch, Mattes, et al.
Published: (2024)
Closed-Form Successive Relative Transfer Function Vector Estimation based on Blind Oblique Projection Incorporating Noise Whitening
by: Gode, Henri, et al.
Published: (2025)
by: Gode, Henri, et al.
Published: (2025)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
Low-Complexity Own Voice Reconstruction for Hearables with an In-Ear Microphone
by: Ohlenbusch, Mattes, et al.
Published: (2024)
by: Ohlenbusch, Mattes, et al.
Published: (2024)
SALAD-VAE: Semantic Audio Compression with Language-Audio Distillation
by: Braun, Sebastian, et al.
Published: (2025)
by: Braun, Sebastian, et al.
Published: (2025)
Modeling of Speech-dependent Own Voice Transfer Characteristics for Hearables with In-ear Microphones
by: Ohlenbusch, Mattes, et al.
Published: (2023)
by: Ohlenbusch, Mattes, et al.
Published: (2023)
Speech-dependent Modeling of Own Voice Transfer Characteristics for In-ear Microphones in Hearables
by: Ohlenbusch, Mattes, et al.
Published: (2023)
by: Ohlenbusch, Mattes, et al.
Published: (2023)
Coherence-Based Frequency Subset Selection For Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers
by: Fejgin, Daniel, et al.
Published: (2022)
by: Fejgin, Daniel, et al.
Published: (2022)
Does Single-channel Speech Enhancement Improve Keyword Spotting Accuracy? A Case Study
by: Brueggeman, Avamarie, et al.
Published: (2023)
by: Brueggeman, Avamarie, et al.
Published: (2023)
Multi-Microphone Noise Data Augmentation for DNN-based Own Voice Reconstruction for Hearables in Noisy Environments
by: Ohlenbusch, Mattes, et al.
Published: (2023)
by: Ohlenbusch, Mattes, et al.
Published: (2023)
VAE-based Phoneme Alignment Using Gradient Annealing and SSL Acoustic Features
by: Koriyama, Tomoki
Published: (2024)
by: Koriyama, Tomoki
Published: (2024)
VQalAttent: a Transparent Speech Generation Pipeline based on Transformer-learned VQ-VAE Latent Space
by: Rodriguez, Armani, et al.
Published: (2024)
by: Rodriguez, Armani, et al.
Published: (2024)
Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers
by: Fejgin, Daniel, et al.
Published: (2025)
by: Fejgin, Daniel, et al.
Published: (2025)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
by: Melechovsky, Jan, et al.
Published: (2024)
by: Melechovsky, Jan, et al.
Published: (2024)
Multi-Source Position and Direction-of-Arrival Estimation Based on Euclidean Distance Matrices
by: Brümann, Klaus, et al.
Published: (2025)
by: Brümann, Klaus, et al.
Published: (2025)
DNN-Based Online Source Counting Based on Spatial Generalized Magnitude Squared Coherence
by: Gode, Henri, et al.
Published: (2026)
by: Gode, Henri, et al.
Published: (2026)
Assisted RTF-Vector-Based Binaural Direction of Arrival Estimation Exploiting a Calibrated External Microphone Array
by: Fejgin, Daniel, et al.
Published: (2022)
by: Fejgin, Daniel, et al.
Published: (2022)
Unrestricted Global Phase Bias-Aware Single-channel Speech Enhancement with Conformer-based Metric GAN
by: Zhang, Shiqi, et al.
Published: (2024)
by: Zhang, Shiqi, et al.
Published: (2024)
Incremental Averaging Method to Improve Graph-Based Time-Difference-of-Arrival Estimation
by: Brümann, Klaus, et al.
Published: (2025)
by: Brümann, Klaus, et al.
Published: (2025)
Learning Time-Graph Frequency Representation for Monaural Speech Enhancement
by: Wang, Tingting, et al.
Published: (2025)
by: Wang, Tingting, et al.
Published: (2025)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
by: Ma, Ding, et al.
Published: (2026)
by: Ma, Ding, et al.
Published: (2026)
Improved Topology-Independent Distributed Adaptive Node-Specific Signal Estimation for Wireless Acoustic Sensor Networks
by: Didier, Paul, et al.
Published: (2025)
by: Didier, Paul, et al.
Published: (2025)
Steered Response Power-Based Direction-of-Arrival Estimation Exploiting an Auxiliary Microphone
by: Brümann, Klaus, et al.
Published: (2024)
by: Brümann, Klaus, et al.
Published: (2024)
Exploiting an External Microphone for Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers
by: Fejgin, Daniel, et al.
Published: (2023)
by: Fejgin, Daniel, et al.
Published: (2023)
EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding
by: Cerovaz, Luca, et al.
Published: (2026)
by: Cerovaz, Luca, et al.
Published: (2026)
Effect of target signals and delays on spatially selective active noise control for open-fitting hearables
by: Xiao, Tong, et al.
Published: (2024)
by: Xiao, Tong, et al.
Published: (2024)
METEOR: Melody-aware Texture-controllable Symbolic Orchestral Music Generation via Transformer VAE
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)
by: Le, Dinh-Viet-Toan, et al.
Published: (2024)
Head Orientation Estimation with Distributed Microphones Using Speech Radiation Patterns
by: Müller, Kaspar, et al.
Published: (2023)
by: Müller, Kaspar, et al.
Published: (2023)
Binaural Localization Model for Speech in Noise
by: Tokala, Vikas, et al.
Published: (2025)
by: Tokala, Vikas, et al.
Published: (2025)
Complex Recurrent Variational Autoencoder with Application to Speech Enhancement
by: Xie, Yuying, et al.
Published: (2022)
by: Xie, Yuying, et al.
Published: (2022)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
by: Wang, Shih-heng, et al.
Published: (2024)
by: Wang, Shih-heng, et al.
Published: (2024)
Similar Items
-
Investigation of Speech and Noise Latent Representations in Single-channel VAE-based Speech Enhancement
by: Li, Jiatong, et al.
Published: (2025) -
Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
by: Han, Runduo, et al.
Published: (2024) -
Binaural Speech Enhancement Using Deep Complex Convolutional Transformer Networks
by: Tokala, Vikas, et al.
Published: (2024) -
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
by: Niu, Zhikang, et al.
Published: (2025) -
Flexible Multi-Channel Target Speaker Extraction Using Geometry-Conditioned Spatially Selective Non-linear Filters
by: Li, Jiatong, et al.
Published: (2026)