DualSep: A Light-weight dual-encoder convolutional recurrent network for real-time in-car speech separation
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Ziqian, Sun, Jiayao, Zhang, Zihan, Li, Xingchen, Liu, Jie, Xie, Lei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
di: Zhu, Yike, et al.
Pubblicazione: (2025)
di: Zhu, Yike, et al.
Pubblicazione: (2025)
Distilling a speech and music encoder with task arithmetic
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
WhisperFlow: speech foundation models in real time
di: Wang, Rongxiang, et al.
Pubblicazione: (2024)
di: Wang, Rongxiang, et al.
Pubblicazione: (2024)
Omni-directional attention mechanism based on Mamba for speech separation
di: Xue, Ke, et al.
Pubblicazione: (2026)
di: Xue, Ke, et al.
Pubblicazione: (2026)
SPGM: Prioritizing Local Features for enhanced speech separation performance
di: Yip, Jia Qi, et al.
Pubblicazione: (2023)
di: Yip, Jia Qi, et al.
Pubblicazione: (2023)
DualVC 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion
di: Ning, Ziqian, et al.
Pubblicazione: (2023)
di: Ning, Ziqian, et al.
Pubblicazione: (2023)
Multichannel blind speech source separation with a disjoint constraint source model
di: Wang, Jianyu, et al.
Pubblicazione: (2024)
di: Wang, Jianyu, et al.
Pubblicazione: (2024)
Acoustic characterization of speech rhythm: going beyond metrics with recurrent neural networks
di: Deloche, François, et al.
Pubblicazione: (2024)
di: Deloche, François, et al.
Pubblicazione: (2024)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
di: Guo, Zhao, et al.
Pubblicazione: (2025)
di: Guo, Zhao, et al.
Pubblicazione: (2025)
A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
di: Laperrière, Gaëlle, et al.
Pubblicazione: (2024)
di: Laperrière, Gaëlle, et al.
Pubblicazione: (2024)
An adaptive filter bank based neural network approach for time delay estimation and speech enhancement
di: Ma, Lu
Pubblicazione: (2025)
di: Ma, Lu
Pubblicazione: (2025)
Determined blind source separation via modeling adjacent frequency band correlations in speech signals
di: Wang, Jianyu, et al.
Pubblicazione: (2025)
di: Wang, Jianyu, et al.
Pubblicazione: (2025)
Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
di: Ma, Linhan, et al.
Pubblicazione: (2024)
di: Ma, Linhan, et al.
Pubblicazione: (2024)
Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
di: Ning, Ziqian, et al.
Pubblicazione: (2024)
di: Ning, Ziqian, et al.
Pubblicazione: (2024)
RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
di: Liu, Mingshuai, et al.
Pubblicazione: (2024)
di: Liu, Mingshuai, et al.
Pubblicazione: (2024)
Hybrid-Sep: Language-queried audio source separation via pre-trained Model Fusion and Adversarial Diffusion Training
di: Feng, Jianyuan, et al.
Pubblicazione: (2025)
di: Feng, Jianyuan, et al.
Pubblicazione: (2025)
Zipformer: A faster and better encoder for automatic speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2023)
di: Yao, Zengwei, et al.
Pubblicazione: (2023)
A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
di: Serre, Thomas, et al.
Pubblicazione: (2024)
di: Serre, Thomas, et al.
Pubblicazione: (2024)
Towards detecting the pathological subharmonic voicing with fully convolutional neural networks
di: Ikuma, Takeshi, et al.
Pubblicazione: (2025)
di: Ikuma, Takeshi, et al.
Pubblicazione: (2025)
Enhancement by postfiltering for speech and audio coding in ad-hoc sensor networks
di: Das, Sneha, et al.
Pubblicazione: (2020)
di: Das, Sneha, et al.
Pubblicazione: (2020)
A correlation-permutation approach for speech-music encoders model merging
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
Accent-VITS:accent transfer for end-to-end TTS
di: Ma, Linhan, et al.
Pubblicazione: (2023)
di: Ma, Linhan, et al.
Pubblicazione: (2023)
Multispecies bird sound recognition using a fully convolutional neural network
di: García-Ordás, María Teresa, et al.
Pubblicazione: (2024)
di: García-Ordás, María Teresa, et al.
Pubblicazione: (2024)
Graph-based multi-Feature fusion method for speech emotion recognition
di: Liu, Xueyu, et al.
Pubblicazione: (2024)
di: Liu, Xueyu, et al.
Pubblicazione: (2024)
BFA: Real-time Multilingual Text-to-speech Forced Alignment
di: Rehman, Abdul, et al.
Pubblicazione: (2025)
di: Rehman, Abdul, et al.
Pubblicazione: (2025)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
di: Dai, Yuhang, et al.
Pubblicazione: (2025)
di: Dai, Yuhang, et al.
Pubblicazione: (2025)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
di: Okocha, Chibuzor, et al.
Pubblicazione: (2025)
di: Okocha, Chibuzor, et al.
Pubblicazione: (2025)
SepMamba: State-space models for speaker separation using Mamba
di: Avenstrup, Thor Højhus, et al.
Pubblicazione: (2024)
di: Avenstrup, Thor Højhus, et al.
Pubblicazione: (2024)
Adaptive high-precision sound source localization at low frequencies based on convolutional neural network
di: Ma, Wenbo, et al.
Pubblicazione: (2024)
di: Ma, Wenbo, et al.
Pubblicazione: (2024)
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
di: Wang, Yi, et al.
Pubblicazione: (2025)
di: Wang, Yi, et al.
Pubblicazione: (2025)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
di: Yao, Jixun, et al.
Pubblicazione: (2024)
di: Yao, Jixun, et al.
Pubblicazione: (2024)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
di: Ma, Guobin, et al.
Pubblicazione: (2025)
di: Ma, Guobin, et al.
Pubblicazione: (2025)
What do neural networks listen to? Exploring the crucial bands in Speech Enhancement using Sinc-convolution
di: Ho, Kuan-Hsun, et al.
Pubblicazione: (2024)
di: Ho, Kuan-Hsun, et al.
Pubblicazione: (2024)
Adversarial speech for voice privacy protection from Personalized Speech generation
di: Chen, Shihao, et al.
Pubblicazione: (2024)
di: Chen, Shihao, et al.
Pubblicazione: (2024)
SPMamba: State-space model is all you need in speech separation
di: Li, Kai, et al.
Pubblicazione: (2024)
di: Li, Kai, et al.
Pubblicazione: (2024)
KS-Net: Multi-band joint speech restoration and enhancement network for 2024 ICASSP SSI Challenge
di: Yu, Guochen, et al.
Pubblicazione: (2024)
di: Yu, Guochen, et al.
Pubblicazione: (2024)
U-Mamba-Net: A highly efficient Mamba-based U-net style network for noisy and reverberant speech separation
di: Dang, Shaoxiang, et al.
Pubblicazione: (2024)
di: Dang, Shaoxiang, et al.
Pubblicazione: (2024)
Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
di: Han, Runduo, et al.
Pubblicazione: (2024)
di: Han, Runduo, et al.
Pubblicazione: (2024)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
di: Zhu, Yike, et al.
Pubblicazione: (2025) -
Distilling a speech and music encoder with task arithmetic
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025) -
BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators
di: Zhang, Zihan, et al.
Pubblicazione: (2024) -
WhisperFlow: speech foundation models in real time
di: Wang, Rongxiang, et al.
Pubblicazione: (2024) -
Omni-directional attention mechanism based on Mamba for speech separation
di: Xue, Ke, et al.
Pubblicazione: (2026)