Long-form music generation with latent diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Evans, Zach, Parker, Julian D., Carr, CJ, Zukowski, Zack, Taylor, Josiah, Pons, Jordi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stable Audio Open
by: Evans, Zach, et al.
Published: (2024)
by: Evans, Zach, et al.
Published: (2024)
Music and Artificial Intelligence: Artistic Trends
by: Pons, Jordi, et al.
Published: (2025)
by: Pons, Jordi, et al.
Published: (2025)
Fast Timing-Conditioned Latent Audio Diffusion
by: Evans, Zach, et al.
Published: (2024)
by: Evans, Zach, et al.
Published: (2024)
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
by: Parker, Julian D, et al.
Published: (2024)
by: Parker, Julian D, et al.
Published: (2024)
Fast Text-to-Audio Generation with Adversarial Post-Training
by: Novack, Zachary, et al.
Published: (2025)
by: Novack, Zachary, et al.
Published: (2025)
StemGen: A music generation model that listens
by: Parker, Julian D., et al.
Published: (2023)
by: Parker, Julian D., et al.
Published: (2023)
Combining audio control and style transfer using latent diffusion
by: Demerlé, Nils, et al.
Published: (2024)
by: Demerlé, Nils, et al.
Published: (2024)
SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
by: Ahmed, Tawsif, et al.
Published: (2025)
by: Ahmed, Tawsif, et al.
Published: (2025)
Approaching an unknown communication system by latent space exploration and causal inference
by: Beguš, Gašper, et al.
Published: (2023)
by: Beguš, Gašper, et al.
Published: (2023)
Detecting music deepfakes is easy but actually hard
by: Afchar, Darius, et al.
Published: (2024)
by: Afchar, Darius, et al.
Published: (2024)
Computational music analysis from first principles
by: Tymoczko, Dmitri, et al.
Published: (2024)
by: Tymoczko, Dmitri, et al.
Published: (2024)
Low-Resource Guidance for Controllable Latent Audio Diffusion
by: Novack, Zachary, et al.
Published: (2026)
by: Novack, Zachary, et al.
Published: (2026)
Linear RNNs for autoregressive generation of long music samples
by: Szewczyk, Konrad, et al.
Published: (2025)
by: Szewczyk, Konrad, et al.
Published: (2025)
Symbotunes: unified hub for symbolic music generative models
by: Skierś, Paweł, et al.
Published: (2024)
by: Skierś, Paweł, et al.
Published: (2024)
Learning and composing of classical music using restricted Boltzmann machines
by: Kobayashi, Mutsumi, et al.
Published: (2025)
by: Kobayashi, Mutsumi, et al.
Published: (2025)
Local deployment of large-scale music AI models on commodity hardware
by: Zhou, Xun, et al.
Published: (2024)
by: Zhou, Xun, et al.
Published: (2024)
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
by: Messina, Francisco, et al.
Published: (2025)
by: Messina, Francisco, et al.
Published: (2025)
The first Cadenza challenges: using machine learning competitions to improve music for listeners with a hearing loss
by: Dabike, Gerardo Roa, et al.
Published: (2024)
by: Dabike, Gerardo Roa, et al.
Published: (2024)
Female mosquito detection by means of AI techniques inside release containers in the context of a Sterile Insect Technique program
by: Naranjo-Alcazar, Javier, et al.
Published: (2023)
by: Naranjo-Alcazar, Javier, et al.
Published: (2023)
Context-aware child-directed speech detection from long-form recordings
by: Charlot, Théo, et al.
Published: (2026)
by: Charlot, Théo, et al.
Published: (2026)
Testing chatbots on the creation of encoders for audio conditioned image generation
by: León, Jorge E., et al.
Published: (2025)
by: León, Jorge E., et al.
Published: (2025)
PerTok: Expressive Encoding and Modeling of Symbolic Musical Ideas and Variations
by: Lenz, Julian, et al.
Published: (2024)
by: Lenz, Julian, et al.
Published: (2024)
A Data-Centric Framework for Machine Listening Projects: Addressing Large-Scale Data Acquisition and Labeling through Active Learning
by: Naranjo-Alcazar, Javier, et al.
Published: (2024)
by: Naranjo-Alcazar, Javier, et al.
Published: (2024)
Improving AI-generated music with user-guided training
by: Singh, Vishwa Mohan, et al.
Published: (2025)
by: Singh, Vishwa Mohan, et al.
Published: (2025)
Similarity-Guided Diffusion for Long-Gap Music Inpainting
by: Turland, Sean, et al.
Published: (2025)
by: Turland, Sean, et al.
Published: (2025)
BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music
by: Yao, Mingyang, et al.
Published: (2025)
by: Yao, Mingyang, et al.
Published: (2025)
An Attention Long Short-Term Memory based system for automatic classification of speech intelligibility
by: Fernández-Díaz, Miguel, et al.
Published: (2024)
by: Fernández-Díaz, Miguel, et al.
Published: (2024)
Integrating Text-to-Music Models with Language Models: Composing Long Structured Music Pieces
by: Atassi, Lilac
Published: (2024)
by: Atassi, Lilac
Published: (2024)
Fifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to Validity
by: Peurey, Loann, et al.
Published: (2025)
by: Peurey, Loann, et al.
Published: (2025)
BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings
by: Charlot, Théo, et al.
Published: (2025)
by: Charlot, Théo, et al.
Published: (2025)
Stable Audio 3
by: Evans, Zach, et al.
Published: (2026)
by: Evans, Zach, et al.
Published: (2026)
Zero-Shot Mono-to-Binaural Speech Synthesis
by: Levkovitch, Alon, et al.
Published: (2024)
by: Levkovitch, Alon, et al.
Published: (2024)
STASE: A spatialized text-to-audio synthesis engine for music generation
by: Chi, Tutti, et al.
Published: (2025)
by: Chi, Tutti, et al.
Published: (2025)
SoundReactor: Frame-level Online Video-to-Audio Generation
by: Saito, Koichi, et al.
Published: (2025)
by: Saito, Koichi, et al.
Published: (2025)
The 2025 PNPL Competition: Speech Detection and Phoneme Classification in the LibriBrain Dataset
by: Landau, Gilad, et al.
Published: (2025)
by: Landau, Gilad, et al.
Published: (2025)
Avoiding an AI-imposed Taylor's Version of all music history
by: Collins, Nick, et al.
Published: (2024)
by: Collins, Nick, et al.
Published: (2024)
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
by: Rouard, Simon, et al.
Published: (2025)
by: Rouard, Simon, et al.
Published: (2025)
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
by: Chen, Tuochao, et al.
Published: (2025)
by: Chen, Tuochao, et al.
Published: (2025)
Modulation Discovery with Differentiable Digital Signal Processing
by: Mitcheltree, Christopher, et al.
Published: (2025)
by: Mitcheltree, Christopher, et al.
Published: (2025)
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
Similar Items
-
Stable Audio Open
by: Evans, Zach, et al.
Published: (2024) -
Music and Artificial Intelligence: Artistic Trends
by: Pons, Jordi, et al.
Published: (2025) -
Fast Timing-Conditioned Latent Audio Diffusion
by: Evans, Zach, et al.
Published: (2024) -
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
by: Parker, Julian D, et al.
Published: (2024) -
Fast Text-to-Audio Generation with Adversarial Post-Training
by: Novack, Zachary, et al.
Published: (2025)