Adaptable Symbolic Music Infilling with MIDI-RWKV
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou-Zheng, Christian, Pasquier, Philippe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
por: Dhiman, Jai
Publicado: (2026)
por: Dhiman, Jai
Publicado: (2026)
Quantum-Enhanced Analysis and Grading of Vocal Performance
por: Agarwal, Rohan
Publicado: (2025)
por: Agarwal, Rohan
Publicado: (2025)
Prevailing Research Areas for Music AI in the Era of Foundation Models
por: Wei, Megan, et al.
Publicado: (2024)
por: Wei, Megan, et al.
Publicado: (2024)
PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation
por: Yi, Yungang, et al.
Publicado: (2024)
por: Yi, Yungang, et al.
Publicado: (2024)
Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering
por: Aristorenas, Aris J.
Publicado: (2024)
por: Aristorenas, Aris J.
Publicado: (2024)
Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network
por: He, Zhanhong, et al.
Publicado: (2025)
por: He, Zhanhong, et al.
Publicado: (2025)
GraFPrint: A GNN-Based Approach for Audio Identification
por: Bhattacharjee, Aditya, et al.
Publicado: (2024)
por: Bhattacharjee, Aditya, et al.
Publicado: (2024)
Scalable Evaluation for Audio Identification via Synthetic Latent Fingerprint Generation
por: Bhattacharjee, Aditya, et al.
Publicado: (2025)
por: Bhattacharjee, Aditya, et al.
Publicado: (2025)
Automatic Album Sequencing
por: Herrmann, Vincent, et al.
Publicado: (2024)
por: Herrmann, Vincent, et al.
Publicado: (2024)
A Multimodal Symphony: Integrating Taste and Sound through Generative AI
por: Spanio, Matteo, et al.
Publicado: (2025)
por: Spanio, Matteo, et al.
Publicado: (2025)
Score Distillation Sampling for Audio: Source Separation, Synthesis, and Beyond
por: Richter-Powell, Jessie, et al.
Publicado: (2025)
por: Richter-Powell, Jessie, et al.
Publicado: (2025)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
por: Mehta, Shivam, et al.
Publicado: (2025)
por: Mehta, Shivam, et al.
Publicado: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
por: Mehta, Shivam, et al.
Publicado: (2025)
por: Mehta, Shivam, et al.
Publicado: (2025)
Generation of Musical Timbres using a Text-Guided Diffusion Model
por: Yuan, Weixuan, et al.
Publicado: (2025)
por: Yuan, Weixuan, et al.
Publicado: (2025)
DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids
por: Tsangko, Iosif, et al.
Publicado: (2025)
por: Tsangko, Iosif, et al.
Publicado: (2025)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
por: Mehta, Shivam, et al.
Publicado: (2024)
por: Mehta, Shivam, et al.
Publicado: (2024)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
por: Kim, Minu, et al.
Publicado: (2025)
por: Kim, Minu, et al.
Publicado: (2025)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
por: Wang, Shaowen, et al.
Publicado: (2025)
por: Wang, Shaowen, et al.
Publicado: (2025)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
por: Li, Pengcheng, et al.
Publicado: (2024)
por: Li, Pengcheng, et al.
Publicado: (2024)
OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression
por: Mahfi, Muntahi Safwan, et al.
Publicado: (2025)
por: Mahfi, Muntahi Safwan, et al.
Publicado: (2025)
MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation
por: Zhu, Di, et al.
Publicado: (2026)
por: Zhu, Di, et al.
Publicado: (2026)
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
por: Kim, Minu, et al.
Publicado: (2025)
por: Kim, Minu, et al.
Publicado: (2025)
Benchmarking Sub-Genre Classification For Mainstage Dance Music
por: Shu, Hongzhi, et al.
Publicado: (2024)
por: Shu, Hongzhi, et al.
Publicado: (2024)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
por: Khushiyant, et al.
Publicado: (2026)
por: Khushiyant, et al.
Publicado: (2026)
BemaGANv2: Discriminator Combination Strategies for GAN-based Vocoders in Long-Term Audio Generation
por: Park, Taesoo, et al.
Publicado: (2025)
por: Park, Taesoo, et al.
Publicado: (2025)
Matcha-TTS: A fast TTS architecture with conditional flow matching
por: Mehta, Shivam, et al.
Publicado: (2023)
por: Mehta, Shivam, et al.
Publicado: (2023)
Hidden Echoes Survive Training in Audio To Audio Generative Instrument Models
por: Tralie, Christopher J., et al.
Publicado: (2024)
por: Tralie, Christopher J., et al.
Publicado: (2024)
Listen to the Unexpected: Self-Supervised Surprise Detection for Efficient Viewport Prediction
por: Khah, Arman Nik, et al.
Publicado: (2026)
por: Khah, Arman Nik, et al.
Publicado: (2026)
Modeling L1 Influence on L2 Pronunciation: An MFCC-Based Framework for Explainable Machine Learning and Pedagogical Feedback
por: Jahanbin, Peyman
Publicado: (2025)
por: Jahanbin, Peyman
Publicado: (2025)
M6(GPT)3: Generating Multitrack Modifiable Multi-Minute MIDI Music from Text using Genetic algorithms, Probabilistic methods and GPT Models in any Progression and Time Signature
por: Poćwiardowski, Jakub, et al.
Publicado: (2024)
por: Poćwiardowski, Jakub, et al.
Publicado: (2024)
acoupi: An Open-Source Python Framework for Deploying Bioacoustic AI Models on Edge Devices
por: Vuilliomenet, Aude, et al.
Publicado: (2025)
por: Vuilliomenet, Aude, et al.
Publicado: (2025)
Revisiting SSL for sound event detection: complementary fusion and adaptive post-processing
por: Cui, Hanfang, et al.
Publicado: (2025)
por: Cui, Hanfang, et al.
Publicado: (2025)
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
por: Melechovsky, Jan, et al.
Publicado: (2025)
por: Melechovsky, Jan, et al.
Publicado: (2025)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
por: Chen, Kuan-Yu, et al.
Publicado: (2025)
por: Chen, Kuan-Yu, et al.
Publicado: (2025)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
por: Donepudi, Dharma Teja
Publicado: (2025)
por: Donepudi, Dharma Teja
Publicado: (2025)
Taming Audio VAEs via Target-KL Regularization
por: Seetharaman, Prem, et al.
Publicado: (2026)
por: Seetharaman, Prem, et al.
Publicado: (2026)
Fine-tuning Pre-trained Audio Models for COVID-19 Detection: A Technical Report
por: de Brito, Daniel Oliveira, et al.
Publicado: (2025)
por: de Brito, Daniel Oliveira, et al.
Publicado: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
por: Viswanathan, Janaki, et al.
Publicado: (2025)
por: Viswanathan, Janaki, et al.
Publicado: (2025)
Reciprocal Latent Fields for Precomputed Sound Propagation
por: Seuté, Hugo, et al.
Publicado: (2026)
por: Seuté, Hugo, et al.
Publicado: (2026)
PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
por: Vosoughi, Ali, et al.
Publicado: (2025)
por: Vosoughi, Ali, et al.
Publicado: (2025)
Ejemplares similares
-
Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation
por: Dhiman, Jai
Publicado: (2026) -
Quantum-Enhanced Analysis and Grading of Vocal Performance
por: Agarwal, Rohan
Publicado: (2025) -
Prevailing Research Areas for Music AI in the Era of Foundation Models
por: Wei, Megan, et al.
Publicado: (2024) -
PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation
por: Yi, Yungang, et al.
Publicado: (2024) -
Machine Learning Framework for Audio-Based Content Evaluation using MFCC, Chroma, Spectral Contrast, and Temporal Feature Engineering
por: Aristorenas, Aris J.
Publicado: (2024)