JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Renhang, Hung, Chia-Yu, Majumder, Navonil, Gautreaux, Taylor, Bagherzadeh, Amir Ali, Li, Chuan, Herremans, Dorien, Poria, Soujanya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model
von: Kang, Jaeyong, et al.
Veröffentlicht: (2023)
von: Kang, Jaeyong, et al.
Veröffentlicht: (2023)
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
von: Liu, Renhang, et al.
Veröffentlicht: (2024)
von: Liu, Renhang, et al.
Veröffentlicht: (2024)
JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata
von: Roy, Abhinaba, et al.
Veröffentlicht: (2025)
von: Roy, Abhinaba, et al.
Veröffentlicht: (2025)
NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2025)
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2025)
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
von: Majumder, Navonil, et al.
Veröffentlicht: (2024)
von: Majumder, Navonil, et al.
Veröffentlicht: (2024)
Mustango: Toward Controllable Text-to-Music Generation
von: Melechovsky, Jan, et al.
Veröffentlicht: (2023)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2023)
Inference Time Alignment with Reward-Guided Tree Search
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
MERIT: Learning Disentangled Music Representations for Audio Similarity
von: Roy, Abhinaba, et al.
Veröffentlicht: (2026)
von: Roy, Abhinaba, et al.
Veröffentlicht: (2026)
Are We There Yet? A Brief Survey of Music Emotion Prediction Datasets, Models and Outstanding Challenges
von: Kang, Jaeyong, et al.
Veröffentlicht: (2024)
von: Kang, Jaeyong, et al.
Veröffentlicht: (2024)
Towards Unified Music Emotion Recognition across Dimensional and Categorical Models
von: Kang, Jaeyong, et al.
Veröffentlicht: (2025)
von: Kang, Jaeyong, et al.
Veröffentlicht: (2025)
Aligning Generative Music AI with Human Preferences: Methods and Challenges
von: Herremans, Dorien, et al.
Veröffentlicht: (2025)
von: Herremans, Dorien, et al.
Veröffentlicht: (2025)
Improving Text-To-Audio Models with Synthetic Captions
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment
von: Roy, Abhinaba, et al.
Veröffentlicht: (2025)
von: Roy, Abhinaba, et al.
Veröffentlicht: (2025)
APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music
von: Husain, Jaavid Aktar, et al.
Veröffentlicht: (2026)
von: Husain, Jaavid Aktar, et al.
Veröffentlicht: (2026)
10 Open Challenges Steering the Future of Vision-Language-Action Models
von: Poria, Soujanya, et al.
Veröffentlicht: (2025)
von: Poria, Soujanya, et al.
Veröffentlicht: (2025)
BandCondiNet: Parallel Transformers-based Conditional Popular Music Generation with Multi-View Features
von: Luo, Jing, et al.
Veröffentlicht: (2024)
von: Luo, Jing, et al.
Veröffentlicht: (2024)
MidiCaps: A large-scale MIDI dataset with text captions
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
MIRFLEX: Music Information Retrieval Feature Library for Extraction
von: Chopra, Anuradha, et al.
Veröffentlicht: (2024)
von: Chopra, Anuradha, et al.
Veröffentlicht: (2024)
SegTune: Structured and Fine-Grained Control for Song Generation
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2025)
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2025)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
von: Ma, Guobin, et al.
Veröffentlicht: (2026)
Text2Score: Generating Sheet Music From Textual Prompts
von: Bhandari, Keshav, et al.
Veröffentlicht: (2026)
von: Bhandari, Keshav, et al.
Veröffentlicht: (2026)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
von: Lam, Perry, et al.
Veröffentlicht: (2022)
von: Lam, Perry, et al.
Veröffentlicht: (2022)
A Machine Learning Approach for MIDI to Guitar Tablature Conversion
von: Kaliakatsos-Papakostas, Maximos, et al.
Veröffentlicht: (2025)
von: Kaliakatsos-Papakostas, Maximos, et al.
Veröffentlicht: (2025)
Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics
von: Song, Maojia, et al.
Veröffentlicht: (2025)
von: Song, Maojia, et al.
Veröffentlicht: (2025)
MelodySim: Measuring Melody-aware Music Similarity for Plagiarism Detection
von: Lu, Tongyu, et al.
Veröffentlicht: (2025)
von: Lu, Tongyu, et al.
Veröffentlicht: (2025)
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
von: Li, Yingting, et al.
Veröffentlicht: (2024)
von: Li, Yingting, et al.
Veröffentlicht: (2024)
Natural Language Processing Methods for Symbolic Music Generation and Information Retrieval: a Survey
von: Le, Dinh-Viet-Toan, et al.
Veröffentlicht: (2024)
von: Le, Dinh-Viet-Toan, et al.
Veröffentlicht: (2024)
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
von: Kwon, Mingi, et al.
Veröffentlicht: (2025)
von: Kwon, Mingi, et al.
Veröffentlicht: (2025)
Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
ImprovNet -- Generating Controllable Musical Improvisations with Iterative Corruption Refinement
von: Bhandari, Keshav, et al.
Veröffentlicht: (2025)
von: Bhandari, Keshav, et al.
Veröffentlicht: (2025)
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
von: Melechovsky, Jan, et al.
Veröffentlicht: (2025)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2025)
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
von: Chopra, Anuradha, et al.
Veröffentlicht: (2025)
von: Chopra, Anuradha, et al.
Veröffentlicht: (2025)
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
von: Li, Yingting, et al.
Veröffentlicht: (2024)
von: Li, Yingting, et al.
Veröffentlicht: (2024)
Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control
von: Jiang, Changhao, et al.
Veröffentlicht: (2026)
von: Jiang, Changhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model
von: Kang, Jaeyong, et al.
Veröffentlicht: (2023) -
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024) -
Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
von: Liu, Renhang, et al.
Veröffentlicht: (2024) -
JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata
von: Roy, Abhinaba, et al.
Veröffentlicht: (2025) -
NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2025)