Saved in:
| Main Authors: | Rika, Daniel, Sapir, Nino, Gus, Ido |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.16420 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
by: Serre, Thomas, et al.
Published: (2024)
by: Serre, Thomas, et al.
Published: (2024)
Automatic Melody Reduction via Shortest Path Finding
by: Wang, Ziyu, et al.
Published: (2025)
by: Wang, Ziyu, et al.
Published: (2025)
Versatile Symbolic Music-for-Music Modeling via Function Alignment
by: Jiang, Junyan, et al.
Published: (2025)
by: Jiang, Junyan, et al.
Published: (2025)
Language Model Mapping in Multimodal Music Learning: A Grand Challenge Proposal
by: Chin, Daniel, et al.
Published: (2025)
by: Chin, Daniel, et al.
Published: (2025)
ViTex: Visual Texture Control for Multi-Track Symbolic Music Generation via Discrete Diffusion Models
by: Yi, Xiaoyu, et al.
Published: (2026)
by: Yi, Xiaoyu, et al.
Published: (2026)
USM RNN-T model weights binarization
by: Rybakov, Oleg, et al.
Published: (2024)
by: Rybakov, Oleg, et al.
Published: (2024)
GhostRNN: Reducing State Redundancy in RNN with Cheap Operations
by: Zhou, Hang, et al.
Published: (2024)
by: Zhou, Hang, et al.
Published: (2024)
Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls
by: Lin, Liwei, et al.
Published: (2024)
by: Lin, Liwei, et al.
Published: (2024)
Structured Multi-Track Accompaniment Arrangement via Style Prior Modelling
by: Zhao, Jingwei, et al.
Published: (2023)
by: Zhao, Jingwei, et al.
Published: (2023)
TOMI: Transforming and Organizing Music Ideas for Multi-Track Compositions with Full-Song Structure
by: He, Qi, et al.
Published: (2025)
by: He, Qi, et al.
Published: (2025)
The Costs of Reproducibility in Music Separation Research: a Replication of Band-Split RNN
by: Magron, Paul, et al.
Published: (2026)
by: Magron, Paul, et al.
Published: (2026)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
by: Yan, Haoyin, et al.
Published: (2024)
by: Yan, Haoyin, et al.
Published: (2024)
M2M-Gen: A Multimodal Framework for Automated Background Music Generation in Japanese Manga Using Large Language Models
by: Sharma, Megha, et al.
Published: (2024)
by: Sharma, Megha, et al.
Published: (2024)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
by: Mattursun, Alimjan, et al.
Published: (2025)
by: Mattursun, Alimjan, et al.
Published: (2025)
CartoonSing: Unifying Human and Nonhuman Timbres in Singing Generation
by: Han, Jionghao, et al.
Published: (2025)
by: Han, Jionghao, et al.
Published: (2025)
Pitch-Aware RNN-T for Mandarin Chinese Mispronunciation Detection and Diagnosis
by: Wang, Xintong, et al.
Published: (2024)
by: Wang, Xintong, et al.
Published: (2024)
Exploring GPT's Ability as a Judge in Music Understanding
by: Fang, Kun, et al.
Published: (2025)
by: Fang, Kun, et al.
Published: (2025)
Content-based Controls For Music Large Language Modeling
by: Lin, Liwei, et al.
Published: (2023)
by: Lin, Liwei, et al.
Published: (2023)
MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
by: Zhao, Shengkui, et al.
Published: (2023)
by: Zhao, Shengkui, et al.
Published: (2023)
Token-Weighted RNN-T for Learning from Flawed Data
by: Keren, Gil, et al.
Published: (2024)
by: Keren, Gil, et al.
Published: (2024)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
by: Le, Khanh, et al.
Published: (2025)
by: Le, Khanh, et al.
Published: (2025)
Two-Path GMM-ResNet and GMM-SENet for ASV Spoofing Detection
by: Lei, Zhenchun, et al.
Published: (2024)
by: Lei, Zhenchun, et al.
Published: (2024)
Assessing the Potential Impact of Direction-Dependent HRTF Selection on Sound Localization Accuracy
by: Goldring, Sapir, et al.
Published: (2024)
by: Goldring, Sapir, et al.
Published: (2024)
Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning
by: Wang, Linge, et al.
Published: (2026)
by: Wang, Linge, et al.
Published: (2026)
Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models
by: Wang, Ziyu, et al.
Published: (2024)
by: Wang, Ziyu, et al.
Published: (2024)
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation
by: Izzati, Fathinah, et al.
Published: (2025)
by: Izzati, Fathinah, et al.
Published: (2025)
The ART of Conversation: Measuring Phonetic Convergence and Deliberate Imitation in L2-Speech with a Siamese RNN
by: Yuan, Zheng, et al.
Published: (2023)
by: Yuan, Zheng, et al.
Published: (2023)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
by: Bataev, Vladimir
Published: (2025)
by: Bataev, Vladimir
Published: (2025)
Serial-Parallel Dual-Path Architecture for Speaking Style Recognition
by: Li, Guojian, et al.
Published: (2025)
by: Li, Guojian, et al.
Published: (2025)
Speech Enhancement with Dual-path Multi-Channel Linear Prediction Filter and Multi-norm Beamforming
by: Qin, Chengyuan, et al.
Published: (2025)
by: Qin, Chengyuan, et al.
Published: (2025)
A Novel CustNetGC Boosted Model with Spectral Features for Parkinson's Disease Prediction
by: Karthik, Abishek, et al.
Published: (2025)
by: Karthik, Abishek, et al.
Published: (2025)
SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection
by: HIdekel, Ido Nitzan, et al.
Published: (2025)
by: HIdekel, Ido Nitzan, et al.
Published: (2025)
RenCon 2025: Revival of the Expressive Performance Rendering Competition
by: Zhang, Huan, et al.
Published: (2026)
by: Zhang, Huan, et al.
Published: (2026)
Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting Networks?
by: Damiano, Stefano, et al.
Published: (2024)
by: Damiano, Stefano, et al.
Published: (2024)
ZipEnhancer: Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement
by: Wang, Haoxu, et al.
Published: (2025)
by: Wang, Haoxu, et al.
Published: (2025)
Dynamic nsNet2: Efficient Deep Noise Suppression with Early Exiting
by: Miccini, Riccardo, et al.
Published: (2023)
by: Miccini, Riccardo, et al.
Published: (2023)
Elastic Net Regularization and Gabor Dictionary for Classification of Heart Sound Signals using Deep Learning
by: Fakhry, Mahmoud, et al.
Published: (2026)
by: Fakhry, Mahmoud, et al.
Published: (2026)
ISAC: An Invertible and Stable Auditory Filter Bank with Customizable Kernels for ML Integration
by: Haider, Daniel, et al.
Published: (2025)
by: Haider, Daniel, et al.
Published: (2025)
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
by: Zhang, Yixiao, et al.
Published: (2024)
by: Zhang, Yixiao, et al.
Published: (2024)
Audio Spoof Detection with GaborNet
by: Maciejko, Waldek
Published: (2026)
by: Maciejko, Waldek
Published: (2026)
Similar Items
-
A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
by: Serre, Thomas, et al.
Published: (2024) -
Automatic Melody Reduction via Shortest Path Finding
by: Wang, Ziyu, et al.
Published: (2025) -
Versatile Symbolic Music-for-Music Modeling via Function Alignment
by: Jiang, Junyan, et al.
Published: (2025) -
Language Model Mapping in Multimodal Music Learning: A Grand Challenge Proposal
by: Chin, Daniel, et al.
Published: (2025) -
ViTex: Visual Texture Control for Multi-Track Symbolic Music Generation via Discrete Diffusion Models
by: Yi, Xiaoyu, et al.
Published: (2026)