Deep learning for music generation. Four approaches and their comparative evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Paroiu, Razvan, Trausan-Matu, Stefan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A correlation-permutation approach for speech-music encoders model merging
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025)
Sound Classification of Four Insect Classes
di: Wang, Yinxuan, et al.
Pubblicazione: (2024)
di: Wang, Yinxuan, et al.
Pubblicazione: (2024)
Echoes: A semantically-aligned music deepfake detection dataset
di: Pascu, Octavian, et al.
Pubblicazione: (2026)
di: Pascu, Octavian, et al.
Pubblicazione: (2026)
Linear RNNs for autoregressive generation of long music samples
di: Szewczyk, Konrad, et al.
Pubblicazione: (2025)
di: Szewczyk, Konrad, et al.
Pubblicazione: (2025)
Symbotunes: unified hub for symbolic music generative models
di: Skierś, Paweł, et al.
Pubblicazione: (2024)
di: Skierś, Paweł, et al.
Pubblicazione: (2024)
InstructAudio: Unified speech and music generation with natural language instruction
di: Qiang, Chunyu, et al.
Pubblicazione: (2025)
di: Qiang, Chunyu, et al.
Pubblicazione: (2025)
Ensemble of classifiers for speech evaluation
di: Belokrylov, G., et al.
Pubblicazione: (2024)
di: Belokrylov, G., et al.
Pubblicazione: (2024)
Automated evaluation of children's speech fluency for low-resource languages
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
DARC: Drum accompaniment generation with fine-grained rhythm control
di: Brosnan, Trey
Pubblicazione: (2026)
di: Brosnan, Trey
Pubblicazione: (2026)
Supervised contrastive learning from weakly-labeled audio segments for musical version matching
di: Serrà, Joan, et al.
Pubblicazione: (2025)
di: Serrà, Joan, et al.
Pubblicazione: (2025)
AFEN: Respiratory Disease Classification using Ensemble Learning
di: Nadkarni, Rahul, et al.
Pubblicazione: (2024)
di: Nadkarni, Rahul, et al.
Pubblicazione: (2024)
Where are we in audio deepfake detection? A systematic analysis over generative and detection models
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
Mode-conditioned music learning and composition: a spiking neural network inspired by neuroscience and psychology
di: Liang, Qian, et al.
Pubblicazione: (2024)
di: Liang, Qian, et al.
Pubblicazione: (2024)
Towards Deep Active Learning in Avian Bioacoustics
di: Rauch, Lukas, et al.
Pubblicazione: (2024)
di: Rauch, Lukas, et al.
Pubblicazione: (2024)
Heart Sound Segmentation Using Deep Learning Techniques
di: Madine, Manas
Pubblicazione: (2024)
di: Madine, Manas
Pubblicazione: (2024)
Arabic Music Classification and Generation using Deep Learning
di: Elshaarawy, Mohamed, et al.
Pubblicazione: (2024)
di: Elshaarawy, Mohamed, et al.
Pubblicazione: (2024)
AND: Audio Network Dissection for Interpreting Deep Acoustic Models
di: Wu, Tung-Yu, et al.
Pubblicazione: (2024)
di: Wu, Tung-Yu, et al.
Pubblicazione: (2024)
Automated data curation for self-supervised learning in underwater acoustic analysis
di: Hummel, Hilde I, et al.
Pubblicazione: (2025)
di: Hummel, Hilde I, et al.
Pubblicazione: (2025)
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
di: Yadav, Sarthak, et al.
Pubblicazione: (2025)
di: Yadav, Sarthak, et al.
Pubblicazione: (2025)
A Hierarchical Deep Learning Approach for Minority Instrument Detection
di: Sechet, Dylan, et al.
Pubblicazione: (2025)
di: Sechet, Dylan, et al.
Pubblicazione: (2025)
SepPrune: Structured Pruning for Efficient Deep Speech Separation
di: Li, Yuqi, et al.
Pubblicazione: (2025)
di: Li, Yuqi, et al.
Pubblicazione: (2025)
Deep Space Separable Distillation for Lightweight Acoustic Scene Classification
di: Ye, ShuQi, et al.
Pubblicazione: (2024)
di: Ye, ShuQi, et al.
Pubblicazione: (2024)
Deep Generic Representations for Domain-Generalized Anomalous Sound Detection
di: Saengthong, Phurich, et al.
Pubblicazione: (2024)
di: Saengthong, Phurich, et al.
Pubblicazione: (2024)
Revisiting Deep Audio-Text Retrieval Through the Lens of Transportation
di: Luong, Manh, et al.
Pubblicazione: (2024)
di: Luong, Manh, et al.
Pubblicazione: (2024)
Sustaining model performance for covid-19 detection from dynamic audio data: Development and evaluation of a comprehensive drift-adaptive framework
di: Ganitidis, Theofanis, et al.
Pubblicazione: (2024)
di: Ganitidis, Theofanis, et al.
Pubblicazione: (2024)
MidiTok Visualizer: a tool for visualization and analysis of tokenized MIDI symbolic music
di: Wiszenko, Michał, et al.
Pubblicazione: (2024)
di: Wiszenko, Michał, et al.
Pubblicazione: (2024)
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
di: Yadav, Sarthak, et al.
Pubblicazione: (2025)
di: Yadav, Sarthak, et al.
Pubblicazione: (2025)
learning discriminative features from spectrograms using center loss for speech emotion recognition
di: Dai, Dongyang, et al.
Pubblicazione: (2025)
di: Dai, Dongyang, et al.
Pubblicazione: (2025)
Effects of Dataset Sampling Rate for Noise Cancellation through Deep Learning
di: Colelough, Brandon, et al.
Pubblicazione: (2024)
di: Colelough, Brandon, et al.
Pubblicazione: (2024)
Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning
di: Guo, Zilu, et al.
Pubblicazione: (2023)
di: Guo, Zilu, et al.
Pubblicazione: (2023)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
Memory-Efficient Training for Deep Speaker Embedding Learning in Speaker Verification
di: Liu, Bei, et al.
Pubblicazione: (2024)
di: Liu, Bei, et al.
Pubblicazione: (2024)
Estimating Musical Surprisal from Audio in Autoregressive Diffusion Model Noise Spaces
di: Bjare, Mathias Rose, et al.
Pubblicazione: (2025)
di: Bjare, Mathias Rose, et al.
Pubblicazione: (2025)
4,500 Seconds: Small Data Training Approaches for Deep UAV Audio Classification
di: Berg, Andrew P., et al.
Pubblicazione: (2025)
di: Berg, Andrew P., et al.
Pubblicazione: (2025)
A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
di: Lu, Shenghui, et al.
Pubblicazione: (2025)
di: Lu, Shenghui, et al.
Pubblicazione: (2025)
Transfer Learning-Based Deep Residual Learning for Speech Recognition in Clean and Noisy Environments
di: Djeffal, Noussaiba, et al.
Pubblicazione: (2025)
di: Djeffal, Noussaiba, et al.
Pubblicazione: (2025)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
di: Pham, Lam, et al.
Pubblicazione: (2024)
di: Pham, Lam, et al.
Pubblicazione: (2024)
Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model
di: Oluwademilade, Adelekun, et al.
Pubblicazione: (2026)
di: Oluwademilade, Adelekun, et al.
Pubblicazione: (2026)
Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
di: Dixit, Satvik, et al.
Pubblicazione: (2024)
di: Dixit, Satvik, et al.
Pubblicazione: (2024)
Audio-to-Image Encoding for Improved Voice Characteristic Detection Using Deep Convolutional Neural Networks
di: Atif, Youness
Pubblicazione: (2025)
di: Atif, Youness
Pubblicazione: (2025)
Documenti analoghi
-
A correlation-permutation approach for speech-music encoders model merging
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2025) -
Sound Classification of Four Insect Classes
di: Wang, Yinxuan, et al.
Pubblicazione: (2024) -
Echoes: A semantically-aligned music deepfake detection dataset
di: Pascu, Octavian, et al.
Pubblicazione: (2026) -
Linear RNNs for autoregressive generation of long music samples
di: Szewczyk, Konrad, et al.
Pubblicazione: (2025) -
Symbotunes: unified hub for symbolic music generative models
di: Skierś, Paweł, et al.
Pubblicazione: (2024)