SMART: Tuning a symbolic music generation system with an audio domain aesthetic reward
Fuente:
arXiv
Guardado en:
| Autores principales: | Jonason, Nicolas, Casini, Luca, Sturm, Bob L. T. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Steer-by-prior Editing of Symbolic Music Loops
por: Jonason, Nicolas, et al.
Publicado: (2024)
por: Jonason, Nicolas, et al.
Publicado: (2024)
SYMPLEX: Controllable Symbolic Music Generation using Simplex Diffusion with Vocabulary Priors
por: Jonason, Nicolas, et al.
Publicado: (2024)
por: Jonason, Nicolas, et al.
Publicado: (2024)
"I made this (sort of)": Negotiating authorship, confronting fraudulence, and exploring new musical spaces with prompt-based AI music generation
por: Sturm, Bob L. T.
Publicado: (2025)
por: Sturm, Bob L. T.
Publicado: (2025)
Data-Driven Analysis of Text-Conditioned AI-Generated Music: A Case Study with Suno and Udio
por: Casini, Luca, et al.
Publicado: (2025)
por: Casini, Luca, et al.
Publicado: (2025)
STASE: A spatialized text-to-audio synthesis engine for music generation
por: Chi, Tutti, et al.
Publicado: (2025)
por: Chi, Tutti, et al.
Publicado: (2025)
Symbotunes: unified hub for symbolic music generative models
por: Skierś, Paweł, et al.
Publicado: (2024)
por: Skierś, Paweł, et al.
Publicado: (2024)
TQCodec: Towards neural audio codec for high-fidelity music streaming
por: He, Lixing, et al.
Publicado: (2026)
por: He, Lixing, et al.
Publicado: (2026)
Audio Dialogues: Dialogues dataset for audio and music understanding
por: Goel, Arushi, et al.
Publicado: (2024)
por: Goel, Arushi, et al.
Publicado: (2024)
Visual-based spatial audio generation system for multi-speaker environments
por: Liu, Xiaojing, et al.
Publicado: (2025)
por: Liu, Xiaojing, et al.
Publicado: (2025)
musif: a Python package for symbolic music feature extraction
por: Llorens, Ana, et al.
Publicado: (2023)
por: Llorens, Ana, et al.
Publicado: (2023)
Training chord recognition models on artificially generated audio
por: Majchrzak, Martyna, et al.
Publicado: (2025)
por: Majchrzak, Martyna, et al.
Publicado: (2025)
Synthetic training set generation using text-to-audio models for environmental sound classification
por: Ronchini, Francesca, et al.
Publicado: (2024)
por: Ronchini, Francesca, et al.
Publicado: (2024)
Joint sentiment analysis of lyrics and audio in music
por: Schaab, Lea, et al.
Publicado: (2024)
por: Schaab, Lea, et al.
Publicado: (2024)
DashengTokenizer: One layer is enough for unified audio understanding and generation
por: Dinkel, Heinrich, et al.
Publicado: (2026)
por: Dinkel, Heinrich, et al.
Publicado: (2026)
Bird detection in audio: a survey and a challenge
por: Stowell, Dan, et al.
Publicado: (2016)
por: Stowell, Dan, et al.
Publicado: (2016)
A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval
por: Oncescu, Andreea-Maria, et al.
Publicado: (2024)
por: Oncescu, Andreea-Maria, et al.
Publicado: (2024)
GLAP: General contrastive audio-text pretraining across domains and languages
por: Dinkel, Heinrich, et al.
Publicado: (2025)
por: Dinkel, Heinrich, et al.
Publicado: (2025)
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
por: Messina, Francisco, et al.
Publicado: (2025)
por: Messina, Francisco, et al.
Publicado: (2025)
Testing chatbots on the creation of encoders for audio conditioned image generation
por: León, Jorge E., et al.
Publicado: (2025)
por: León, Jorge E., et al.
Publicado: (2025)
MidiTok Visualizer: a tool for visualization and analysis of tokenized MIDI symbolic music
por: Wiszenko, Michał, et al.
Publicado: (2024)
por: Wiszenko, Michał, et al.
Publicado: (2024)
TuneGenie: Reasoning-based LLM agents for preferential music generation
por: Pandey, Amitesh, et al.
Publicado: (2025)
por: Pandey, Amitesh, et al.
Publicado: (2025)
Scaling up masked audio encoder learning for general audio classification
por: Dinkel, Heinrich, et al.
Publicado: (2024)
por: Dinkel, Heinrich, et al.
Publicado: (2024)
Recomposer: Event-roll-guided generative audio editing
por: Ellis, Daniel P. W., et al.
Publicado: (2025)
por: Ellis, Daniel P. W., et al.
Publicado: (2025)
vega-mir: An information-theoretic Python toolkit for symbolic music, with applications to harmonic graphs and rubato spectra
por: Jalbert-Desforges, Fred
Publicado: (2026)
por: Jalbert-Desforges, Fred
Publicado: (2026)
Supervised contrastive learning from weakly-labeled audio segments for musical version matching
por: Serrà, Joan, et al.
Publicado: (2025)
por: Serrà, Joan, et al.
Publicado: (2025)
Emoanti: audio anti-deepfake with refined emotion-guided representations
por: Li, Xiaokang, et al.
Publicado: (2025)
por: Li, Xiaokang, et al.
Publicado: (2025)
Making deep neural networks work for medical audio: representation, compression and domain adaptation
por: Onu, Charles C
Publicado: (2025)
por: Onu, Charles C
Publicado: (2025)
An automatic mixing speech enhancement system for multi-track audio
por: Liu, Xiaojing, et al.
Publicado: (2024)
por: Liu, Xiaojing, et al.
Publicado: (2024)
Mixer Metaphors: audio interfaces for non-musical applications
por: McNamara, Tace, et al.
Publicado: (2025)
por: McNamara, Tace, et al.
Publicado: (2025)
AEROMamba: An efficient architecture for audio super-resolution using generative adversarial networks and state space models
por: Abreu, Wallace, et al.
Publicado: (2024)
por: Abreu, Wallace, et al.
Publicado: (2024)
Stage-adaptive audio diffusion modeling
por: Zhang, Xuanhao, et al.
Publicado: (2026)
por: Zhang, Xuanhao, et al.
Publicado: (2026)
GRAM: Spatial general-purpose audio representations for real-world environments
por: Yuksel, Goksenin, et al.
Publicado: (2026)
por: Yuksel, Goksenin, et al.
Publicado: (2026)
A conceptual framework for learning to listen by reward: Curiosity-driven search for novel sources
por: Triantafyllopoulos, Andreas, et al.
Publicado: (2026)
por: Triantafyllopoulos, Andreas, et al.
Publicado: (2026)
A robust audio deepfake detection system via multi-view feature
por: Yang, Yujie, et al.
Publicado: (2024)
por: Yang, Yujie, et al.
Publicado: (2024)
Improved symbolic drum style classification with grammar-based hierarchical representations
por: Géré, Léo, et al.
Publicado: (2024)
por: Géré, Léo, et al.
Publicado: (2024)
A tunable binaural audio telepresence system capable of balancing immersive and enhanced modes
por: Hsu, Yicheng, et al.
Publicado: (2024)
por: Hsu, Yicheng, et al.
Publicado: (2024)
WikiMuTe: A web-sourced dataset of semantic descriptions for music audio
por: Weck, Benno, et al.
Publicado: (2023)
por: Weck, Benno, et al.
Publicado: (2023)
Spatial-CLAP: Learning Spatially-Aware audio--text Embeddings for Multi-Source Conditions
por: Seki, Kentaro, et al.
Publicado: (2025)
por: Seki, Kentaro, et al.
Publicado: (2025)
Long-form music generation with latent diffusion
por: Evans, Zach, et al.
Publicado: (2024)
por: Evans, Zach, et al.
Publicado: (2024)
Towards audio language modeling -- an overview
por: Wu, Haibin, et al.
Publicado: (2024)
por: Wu, Haibin, et al.
Publicado: (2024)
Ejemplares similares
-
Steer-by-prior Editing of Symbolic Music Loops
por: Jonason, Nicolas, et al.
Publicado: (2024) -
SYMPLEX: Controllable Symbolic Music Generation using Simplex Diffusion with Vocabulary Priors
por: Jonason, Nicolas, et al.
Publicado: (2024) -
"I made this (sort of)": Negotiating authorship, confronting fraudulence, and exploring new musical spaces with prompt-based AI music generation
por: Sturm, Bob L. T.
Publicado: (2025) -
Data-Driven Analysis of Text-Conditioned AI-Generated Music: A Case Study with Suno and Udio
por: Casini, Luca, et al.
Publicado: (2025) -
STASE: A spatialized text-to-audio synthesis engine for music generation
por: Chi, Tutti, et al.
Publicado: (2025)