DualSep: A Light-weight dual-encoder convolutional recurrent network for real-time in-car speech separation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Ziqian, Sun, Jiayao, Zhang, Zihan, Li, Xingchen, Liu, Jie, Xie, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
Distilling a speech and music encoder with task arithmetic
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
WhisperFlow: speech foundation models in real time
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
Omni-directional attention mechanism based on Mamba for speech separation
von: Xue, Ke, et al.
Veröffentlicht: (2026)
von: Xue, Ke, et al.
Veröffentlicht: (2026)
SPGM: Prioritizing Local Features for enhanced speech separation performance
von: Yip, Jia Qi, et al.
Veröffentlicht: (2023)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2023)
DualVC 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion
von: Ning, Ziqian, et al.
Veröffentlicht: (2023)
von: Ning, Ziqian, et al.
Veröffentlicht: (2023)
Multichannel blind speech source separation with a disjoint constraint source model
von: Wang, Jianyu, et al.
Veröffentlicht: (2024)
von: Wang, Jianyu, et al.
Veröffentlicht: (2024)
Acoustic characterization of speech rhythm: going beyond metrics with recurrent neural networks
von: Deloche, François, et al.
Veröffentlicht: (2024)
von: Deloche, François, et al.
Veröffentlicht: (2024)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2024)
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2024)
An adaptive filter bank based neural network approach for time delay estimation and speech enhancement
von: Ma, Lu
Veröffentlicht: (2025)
von: Ma, Lu
Veröffentlicht: (2025)
Determined blind source separation via modeling adjacent frequency band correlations in speech signals
von: Wang, Jianyu, et al.
Veröffentlicht: (2025)
von: Wang, Jianyu, et al.
Veröffentlicht: (2025)
Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
von: Ning, Ziqian, et al.
Veröffentlicht: (2024)
von: Ning, Ziqian, et al.
Veröffentlicht: (2024)
RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
Hybrid-Sep: Language-queried audio source separation via pre-trained Model Fusion and Adversarial Diffusion Training
von: Feng, Jianyuan, et al.
Veröffentlicht: (2025)
von: Feng, Jianyuan, et al.
Veröffentlicht: (2025)
Zipformer: A faster and better encoder for automatic speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
von: Serre, Thomas, et al.
Veröffentlicht: (2024)
von: Serre, Thomas, et al.
Veröffentlicht: (2024)
Towards detecting the pathological subharmonic voicing with fully convolutional neural networks
von: Ikuma, Takeshi, et al.
Veröffentlicht: (2025)
von: Ikuma, Takeshi, et al.
Veröffentlicht: (2025)
Enhancement by postfiltering for speech and audio coding in ad-hoc sensor networks
von: Das, Sneha, et al.
Veröffentlicht: (2020)
von: Das, Sneha, et al.
Veröffentlicht: (2020)
A correlation-permutation approach for speech-music encoders model merging
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
Accent-VITS:accent transfer for end-to-end TTS
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
Multispecies bird sound recognition using a fully convolutional neural network
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
Graph-based multi-Feature fusion method for speech emotion recognition
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
von: Liu, Xueyu, et al.
Veröffentlicht: (2024)
BFA: Real-time Multilingual Text-to-speech Forced Alignment
von: Rehman, Abdul, et al.
Veröffentlicht: (2025)
von: Rehman, Abdul, et al.
Veröffentlicht: (2025)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
SepMamba: State-space models for speaker separation using Mamba
von: Avenstrup, Thor Højhus, et al.
Veröffentlicht: (2024)
von: Avenstrup, Thor Højhus, et al.
Veröffentlicht: (2024)
Adaptive high-precision sound source localization at low frequencies based on convolutional neural network
von: Ma, Wenbo, et al.
Veröffentlicht: (2024)
von: Ma, Wenbo, et al.
Veröffentlicht: (2024)
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
What do neural networks listen to? Exploring the crucial bands in Speech Enhancement using Sinc-convolution
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
Adversarial speech for voice privacy protection from Personalized Speech generation
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
SPMamba: State-space model is all you need in speech separation
von: Li, Kai, et al.
Veröffentlicht: (2024)
von: Li, Kai, et al.
Veröffentlicht: (2024)
KS-Net: Multi-band joint speech restoration and enhancement network for 2024 ICASSP SSI Challenge
von: Yu, Guochen, et al.
Veröffentlicht: (2024)
von: Yu, Guochen, et al.
Veröffentlicht: (2024)
U-Mamba-Net: A highly efficient Mamba-based U-net style network for noisy and reverberant speech separation
von: Dang, Shaoxiang, et al.
Veröffentlicht: (2024)
von: Dang, Shaoxiang, et al.
Veröffentlicht: (2024)
Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
von: Han, Runduo, et al.
Veröffentlicht: (2024)
von: Han, Runduo, et al.
Veröffentlicht: (2024)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
von: Tabatabaee, Saba, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025) -
Distilling a speech and music encoder with task arithmetic
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025) -
BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators
von: Zhang, Zihan, et al.
Veröffentlicht: (2024) -
WhisperFlow: speech foundation models in real time
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024) -
Omni-directional attention mechanism based on Mamba for speech separation
von: Xue, Ke, et al.
Veröffentlicht: (2026)