DPDFNet: Boosting DeepFilterNet2 via Dual-Path RNN
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rika, Daniel, Sapir, Nino, Gus, Ido |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
von: Serre, Thomas, et al.
Veröffentlicht: (2024)
von: Serre, Thomas, et al.
Veröffentlicht: (2024)
Automatic Melody Reduction via Shortest Path Finding
von: Wang, Ziyu, et al.
Veröffentlicht: (2025)
von: Wang, Ziyu, et al.
Veröffentlicht: (2025)
Versatile Symbolic Music-for-Music Modeling via Function Alignment
von: Jiang, Junyan, et al.
Veröffentlicht: (2025)
von: Jiang, Junyan, et al.
Veröffentlicht: (2025)
Language Model Mapping in Multimodal Music Learning: A Grand Challenge Proposal
von: Chin, Daniel, et al.
Veröffentlicht: (2025)
von: Chin, Daniel, et al.
Veröffentlicht: (2025)
USM RNN-T model weights binarization
von: Rybakov, Oleg, et al.
Veröffentlicht: (2024)
von: Rybakov, Oleg, et al.
Veröffentlicht: (2024)
ViTex: Visual Texture Control for Multi-Track Symbolic Music Generation via Discrete Diffusion Models
von: Yi, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Yi, Xiaoyu, et al.
Veröffentlicht: (2026)
GhostRNN: Reducing State Redundancy in RNN with Cheap Operations
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
The Costs of Reproducibility in Music Separation Research: a Replication of Band-Split RNN
von: Magron, Paul, et al.
Veröffentlicht: (2026)
von: Magron, Paul, et al.
Veröffentlicht: (2026)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
von: Yan, Haoyin, et al.
Veröffentlicht: (2024)
von: Yan, Haoyin, et al.
Veröffentlicht: (2024)
Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls
von: Lin, Liwei, et al.
Veröffentlicht: (2024)
von: Lin, Liwei, et al.
Veröffentlicht: (2024)
Structured Multi-Track Accompaniment Arrangement via Style Prior Modelling
von: Zhao, Jingwei, et al.
Veröffentlicht: (2023)
von: Zhao, Jingwei, et al.
Veröffentlicht: (2023)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
CartoonSing: Unifying Human and Nonhuman Timbres in Singing Generation
von: Han, Jionghao, et al.
Veröffentlicht: (2025)
von: Han, Jionghao, et al.
Veröffentlicht: (2025)
TOMI: Transforming and Organizing Music Ideas for Multi-Track Compositions with Full-Song Structure
von: He, Qi, et al.
Veröffentlicht: (2025)
von: He, Qi, et al.
Veröffentlicht: (2025)
Pitch-Aware RNN-T for Mandarin Chinese Mispronunciation Detection and Diagnosis
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
von: Wang, Xintong, et al.
Veröffentlicht: (2024)
MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
von: Zhao, Shengkui, et al.
Veröffentlicht: (2023)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2023)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
von: Le, Khanh, et al.
Veröffentlicht: (2025)
von: Le, Khanh, et al.
Veröffentlicht: (2025)
Two-Path GMM-ResNet and GMM-SENet for ASV Spoofing Detection
von: Lei, Zhenchun, et al.
Veröffentlicht: (2024)
von: Lei, Zhenchun, et al.
Veröffentlicht: (2024)
M2M-Gen: A Multimodal Framework for Automated Background Music Generation in Japanese Manga Using Large Language Models
von: Sharma, Megha, et al.
Veröffentlicht: (2024)
von: Sharma, Megha, et al.
Veröffentlicht: (2024)
Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning
von: Wang, Linge, et al.
Veröffentlicht: (2026)
von: Wang, Linge, et al.
Veröffentlicht: (2026)
Token-Weighted RNN-T for Learning from Flawed Data
von: Keren, Gil, et al.
Veröffentlicht: (2024)
von: Keren, Gil, et al.
Veröffentlicht: (2024)
Exploring GPT's Ability as a Judge in Music Understanding
von: Fang, Kun, et al.
Veröffentlicht: (2025)
von: Fang, Kun, et al.
Veröffentlicht: (2025)
Content-based Controls For Music Large Language Modeling
von: Lin, Liwei, et al.
Veröffentlicht: (2023)
von: Lin, Liwei, et al.
Veröffentlicht: (2023)
Assessing the Potential Impact of Direction-Dependent HRTF Selection on Sound Localization Accuracy
von: Goldring, Sapir, et al.
Veröffentlicht: (2024)
von: Goldring, Sapir, et al.
Veröffentlicht: (2024)
Speech Enhancement with Dual-path Multi-Channel Linear Prediction Filter and Multi-norm Beamforming
von: Qin, Chengyuan, et al.
Veröffentlicht: (2025)
von: Qin, Chengyuan, et al.
Veröffentlicht: (2025)
A Novel CustNetGC Boosted Model with Spectral Features for Parkinson's Disease Prediction
von: Karthik, Abishek, et al.
Veröffentlicht: (2025)
von: Karthik, Abishek, et al.
Veröffentlicht: (2025)
Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting Networks?
von: Damiano, Stefano, et al.
Veröffentlicht: (2024)
von: Damiano, Stefano, et al.
Veröffentlicht: (2024)
ZipEnhancer: Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
Serial-Parallel Dual-Path Architecture for Speaking Style Recognition
von: Li, Guojian, et al.
Veröffentlicht: (2025)
von: Li, Guojian, et al.
Veröffentlicht: (2025)
Audio Spoof Detection with GaborNet
von: Maciejko, Waldek
Veröffentlicht: (2026)
von: Maciejko, Waldek
Veröffentlicht: (2026)
TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment
von: Dang, Trung, et al.
Veröffentlicht: (2026)
von: Dang, Trung, et al.
Veröffentlicht: (2026)
The ART of Conversation: Measuring Phonetic Convergence and Deliberate Imitation in L2-Speech with a Siamese RNN
von: Yuan, Zheng, et al.
Veröffentlicht: (2023)
von: Yuan, Zheng, et al.
Veröffentlicht: (2023)
Elastic Net Regularization and Gabor Dictionary for Classification of Heart Sound Signals using Deep Learning
von: Fakhry, Mahmoud, et al.
Veröffentlicht: (2026)
von: Fakhry, Mahmoud, et al.
Veröffentlicht: (2026)
Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models
von: Wang, Ziyu, et al.
Veröffentlicht: (2024)
von: Wang, Ziyu, et al.
Veröffentlicht: (2024)
SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection
von: HIdekel, Ido Nitzan, et al.
Veröffentlicht: (2025)
von: HIdekel, Ido Nitzan, et al.
Veröffentlicht: (2025)
Dynamic nsNet2: Efficient Deep Noise Suppression with Early Exiting
von: Miccini, Riccardo, et al.
Veröffentlicht: (2023)
von: Miccini, Riccardo, et al.
Veröffentlicht: (2023)
RenCon 2025: Revival of the Expressive Performance Rendering Competition
von: Zhang, Huan, et al.
Veröffentlicht: (2026)
von: Zhang, Huan, et al.
Veröffentlicht: (2026)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
von: Bataev, Vladimir
Veröffentlicht: (2025)
von: Bataev, Vladimir
Veröffentlicht: (2025)
ISAC: An Invertible and Stable Auditory Filter Bank with Customizable Kernels for ML Integration
von: Haider, Daniel, et al.
Veröffentlicht: (2025)
von: Haider, Daniel, et al.
Veröffentlicht: (2025)
Are you really listening? Boosting Perceptual Awareness in Music-QA Benchmarks
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
von: Serre, Thomas, et al.
Veröffentlicht: (2024) -
Automatic Melody Reduction via Shortest Path Finding
von: Wang, Ziyu, et al.
Veröffentlicht: (2025) -
Versatile Symbolic Music-for-Music Modeling via Function Alignment
von: Jiang, Junyan, et al.
Veröffentlicht: (2025) -
Language Model Mapping in Multimodal Music Learning: A Grand Challenge Proposal
von: Chin, Daniel, et al.
Veröffentlicht: (2025) -
USM RNN-T model weights binarization
von: Rybakov, Oleg, et al.
Veröffentlicht: (2024)