Local deployment of large-scale music AI models on commodity hardware
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Xun, Ruan, Charlie, Zhao, Zihe, Chen, Tianqi, Donahue, Chris |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
by: Ahmed, Tawsif, et al.
Published: (2025)
by: Ahmed, Tawsif, et al.
Published: (2025)
Anticipatory Music Transformer
by: Thickstun, John, et al.
Published: (2023)
by: Thickstun, John, et al.
Published: (2023)
Just Label the Repeats for In-The-Wild Audio-to-Score Alignment
by: Bukey, Irmak, et al.
Published: (2024)
by: Bukey, Irmak, et al.
Published: (2024)
StemGen: A music generation model that listens
by: Parker, Julian D., et al.
Published: (2023)
by: Parker, Julian D., et al.
Published: (2023)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
by: Long, Phillip, et al.
Published: (2026)
by: Long, Phillip, et al.
Published: (2026)
Detecting music deepfakes is easy but actually hard
by: Afchar, Darius, et al.
Published: (2024)
by: Afchar, Darius, et al.
Published: (2024)
Long-form music generation with latent diffusion
by: Evans, Zach, et al.
Published: (2024)
by: Evans, Zach, et al.
Published: (2024)
Computational music analysis from first principles
by: Tymoczko, Dmitri, et al.
Published: (2024)
by: Tymoczko, Dmitri, et al.
Published: (2024)
Learning and composing of classical music using restricted Boltzmann machines
by: Kobayashi, Mutsumi, et al.
Published: (2025)
by: Kobayashi, Mutsumi, et al.
Published: (2025)
Music Genre Classification: Training an AI model
by: Mogonediwa, Keoikantse
Published: (2024)
by: Mogonediwa, Keoikantse
Published: (2024)
MidiCaps: A large-scale MIDI dataset with text captions
by: Melechovsky, Jan, et al.
Published: (2024)
by: Melechovsky, Jan, et al.
Published: (2024)
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers
by: Dixit, Satvik, et al.
Published: (2024)
by: Dixit, Satvik, et al.
Published: (2024)
Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
by: Lin, Tsung-En, et al.
Published: (2025)
by: Lin, Tsung-En, et al.
Published: (2025)
The first Cadenza challenges: using machine learning competitions to improve music for listeners with a hearing loss
by: Dabike, Gerardo Roa, et al.
Published: (2024)
by: Dabike, Gerardo Roa, et al.
Published: (2024)
Do Music Generation Models Encode Music Theory?
by: Wei, Megan, et al.
Published: (2024)
by: Wei, Megan, et al.
Published: (2024)
wav2pos: Sound Source Localization using Masked Autoencoders
by: Berg, Axel, et al.
Published: (2024)
by: Berg, Axel, et al.
Published: (2024)
Avoiding an AI-imposed Taylor's Version of all music history
by: Collins, Nick, et al.
Published: (2024)
by: Collins, Nick, et al.
Published: (2024)
chatter: a Python library for applying information theory and AI/ML models to animal communication
by: Youngblood, Mason
Published: (2025)
by: Youngblood, Mason
Published: (2025)
An AI-Driven Approach to Wind Turbine Bearing Fault Diagnosis from Acoustic Signals
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
Symbotunes: unified hub for symbolic music generative models
by: Skierś, Paweł, et al.
Published: (2024)
by: Skierś, Paweł, et al.
Published: (2024)
Comparison of spectrogram scaling in multi-label Music Genre Recognition
by: Karpiński, Bartosz, et al.
Published: (2025)
by: Karpiński, Bartosz, et al.
Published: (2025)
Sound Event Detection and Localization with Distance Estimation
by: Krause, Daniel Aleksander, et al.
Published: (2024)
by: Krause, Daniel Aleksander, et al.
Published: (2024)
Audio Simulation for Sound Source Localization in Virtual Evironment
by: Di Yuan, Yi, et al.
Published: (2024)
by: Di Yuan, Yi, et al.
Published: (2024)
Feature Aggregation in Joint Sound Classification and Localization Neural Networks
by: Healy, Brendan, et al.
Published: (2023)
by: Healy, Brendan, et al.
Published: (2023)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
by: Ji, Shengpeng, et al.
Published: (2023)
by: Ji, Shengpeng, et al.
Published: (2023)
Improving Generalization for AI-Synthesized Voice Detection
by: Ren, Hainan, et al.
Published: (2024)
by: Ren, Hainan, et al.
Published: (2024)
Unified AI for Accurate Audio Anomaly Detection
by: Khaleghpour, Hamideh, et al.
Published: (2025)
by: Khaleghpour, Hamideh, et al.
Published: (2025)
An Experimental Study on Joint Modeling for Sound Event Localization and Detection with Source Distance Estimation
by: Dong, Yuxuan, et al.
Published: (2025)
by: Dong, Yuxuan, et al.
Published: (2025)
Mask-Weighted Spatial Likelihood Coding for Speaker-Independent Joint Localization and Mask Estimation
by: Kienegger, Jakob, et al.
Published: (2024)
by: Kienegger, Jakob, et al.
Published: (2024)
Developing an AI-Guided Assistant Device for the Deaf and Hearing Impaired
by: Jiayu, et al.
Published: (2025)
by: Jiayu, et al.
Published: (2025)
Spatial Scaper: A Library to Simulate and Augment Soundscapes for Sound Event Localization and Detection in Realistic Rooms
by: Roman, Iran R., et al.
Published: (2024)
by: Roman, Iran R., et al.
Published: (2024)
Improving AI-generated music with user-guided training
by: Singh, Vishwa Mohan, et al.
Published: (2025)
by: Singh, Vishwa Mohan, et al.
Published: (2025)
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
by: Ji, Shengpeng, et al.
Published: (2024)
by: Ji, Shengpeng, et al.
Published: (2024)
Adversarial Data Augmentation for Robust Speaker Verification
by: Zhou, Zhenyu, et al.
Published: (2024)
by: Zhou, Zhenyu, et al.
Published: (2024)
Uncertainty-Aware Mean Opinion Score Prediction
by: Wang, Hui, et al.
Published: (2024)
by: Wang, Hui, et al.
Published: (2024)
Interpreting Graphic Notation with MusicLDM: An AI Improvisation of Cornelius Cardew's Treatise
by: Karchkhadze, Tornike, et al.
Published: (2024)
by: Karchkhadze, Tornike, et al.
Published: (2024)
On the calibration of powerset speaker diarization models
by: Plaquet, Alexis, et al.
Published: (2024)
by: Plaquet, Alexis, et al.
Published: (2024)
From Sound to Setting: AI-Based Equalizer Parameter Prediction for Piano Tone Replication
by: Yu, Song-Ze
Published: (2025)
by: Yu, Song-Ze
Published: (2025)
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
by: Jiang, Ziyue, et al.
Published: (2025)
by: Jiang, Ziyue, et al.
Published: (2025)
Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers
by: Mittal, Manan, et al.
Published: (2025)
by: Mittal, Manan, et al.
Published: (2025)
Similar Items
-
SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
by: Ahmed, Tawsif, et al.
Published: (2025) -
Anticipatory Music Transformer
by: Thickstun, John, et al.
Published: (2023) -
Just Label the Repeats for In-The-Wild Audio-to-Score Alignment
by: Bukey, Irmak, et al.
Published: (2024) -
StemGen: A music generation model that listens
by: Parker, Julian D., et al.
Published: (2023) -
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
by: Long, Phillip, et al.
Published: (2026)