Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Shuyang, Zhong, Zhi, Wu, Qiyu, Novack, Zachary, Choi, Woosung, Toyama, Keisuke, Cheuk, Kin Wai, Koo, Junghyun, Ikemiya, Yukara, Simon, Christian, Nagashima, Chihiro, Takahashi, Shusuke |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Variable Bitrate Residual Vector Quantization for Audio Coding
by: Chae, Yunkee, et al.
Published: (2024)
by: Chae, Yunkee, et al.
Published: (2024)
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
by: Choi, Woosung, et al.
Published: (2025)
by: Choi, Woosung, et al.
Published: (2025)
Noise-to-Notes: Diffusion-based Generation and Refinement for Automatic Drum Transcription
by: Yeung, Michael, et al.
Published: (2025)
by: Yeung, Michael, et al.
Published: (2025)
Do Foundational Audio Encoders Understand Music Structure?
by: Toyama, Keisuke, et al.
Published: (2025)
by: Toyama, Keisuke, et al.
Published: (2025)
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
by: Comunità, Marco, et al.
Published: (2024)
by: Comunità, Marco, et al.
Published: (2024)
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
by: Zhong, Zhi, et al.
Published: (2025)
by: Zhong, Zhi, et al.
Published: (2025)
Music Foundation Model as Generic Booster for Music Downstream Tasks
by: Liao, WeiHsiang, et al.
Published: (2024)
by: Liao, WeiHsiang, et al.
Published: (2024)
Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription
by: Cwitkowitz, Frank, et al.
Published: (2023)
by: Cwitkowitz, Frank, et al.
Published: (2023)
'Studies for': A Human-AI Co-Creative Sound Artwork Using a Real-time Multi-channel Sound Generation Model
by: Nagashima, Chihiro, et al.
Published: (2025)
by: Nagashima, Chihiro, et al.
Published: (2025)
CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
by: Chen, Yuanhong, et al.
Published: (2025)
by: Chen, Yuanhong, et al.
Published: (2025)
Concept-TRAK: Understanding how diffusion models learn concepts through concept-level attribution
by: Park, Yonghyun, et al.
Published: (2025)
by: Park, Yonghyun, et al.
Published: (2025)
MaskBeat: Loopable Drum Beat Generation
by: Lanzendörfer, Luca A., et al.
Published: (2025)
by: Lanzendörfer, Luca A., et al.
Published: (2025)
DisMix: Disentangling Mixtures of Musical Instruments for Source-level Pitch and Timbre Manipulation
by: Luo, Yin-Jyun, et al.
Published: (2024)
by: Luo, Yin-Jyun, et al.
Published: (2024)
DiffRoll: Diffusion-based Generative Music Transcription with Unsupervised Pretraining Capability
by: Cheuk, Kin Wai, et al.
Published: (2022)
by: Cheuk, Kin Wai, et al.
Published: (2022)
Beat-Based Rhythm Quantization of MIDI Performances
by: Wachter, Maximilian, et al.
Published: (2025)
by: Wachter, Maximilian, et al.
Published: (2025)
LLM2Fx-Tools: Tool Calling For Music Post-Production
by: Doh, Seungheon, et al.
Published: (2025)
by: Doh, Seungheon, et al.
Published: (2025)
Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
by: Zhang, Yixiao, et al.
Published: (2024)
by: Zhang, Yixiao, et al.
Published: (2024)
Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
by: Mancusi, Michele, et al.
Published: (2024)
by: Mancusi, Michele, et al.
Published: (2024)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
by: Long, Phillip, et al.
Published: (2026)
by: Long, Phillip, et al.
Published: (2026)
HQ-VAE: Hierarchical Discrete Representation Learning with Variational Bayes
by: Takida, Yuhta, et al.
Published: (2023)
by: Takida, Yuhta, et al.
Published: (2023)
Automatic Music Mixing using a Generative Model of Effect Embeddings
by: Moliner, Eloi, et al.
Published: (2025)
by: Moliner, Eloi, et al.
Published: (2025)
MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
by: Takahashi, Akira, et al.
Published: (2025)
by: Takahashi, Akira, et al.
Published: (2025)
When the Festival Drums Beat: Demystifying Festival Cuisine in Kerala
by: Asha Krishnan
Published: (2021)
by: Asha Krishnan
Published: (2021)
When the Festival Drums Beat: Demystifying Festival Cuisine in Kerala
by: Asha Krishnan
Published: (2021)
by: Asha Krishnan
Published: (2021)
Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations
by: Wachter, Maximilian, et al.
Published: (2026)
by: Wachter, Maximilian, et al.
Published: (2026)
Beat and Downbeat Tracking in Performance MIDI Using an End-to-End Transformer Architecture
by: Murgul, Sebastian, et al.
Published: (2025)
by: Murgul, Sebastian, et al.
Published: (2025)
Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs
by: Soiledis, Konstantinos, et al.
Published: (2026)
by: Soiledis, Konstantinos, et al.
Published: (2026)
Towards Blind Data Cleaning: A Case Study in Music Source Separation
by: Gui, Azalea, et al.
Published: (2025)
by: Gui, Azalea, et al.
Published: (2025)
A new critical growth parameter and mechanistic model for SiC nanowire synthesis via Si substrate carbonization: the role of H$_2$/CH$_4$ gas flow ratio
by: Koo, Junghyun, et al.
Published: (2024)
by: Koo, Junghyun, et al.
Published: (2024)
MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video
by: Tateishi, Kazuya, et al.
Published: (2026)
by: Tateishi, Kazuya, et al.
Published: (2026)
MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
by: Wu, Shih-Lun, et al.
Published: (2025)
by: Wu, Shih-Lun, et al.
Published: (2025)
Improving Unsupervised Clean-to-Rendered Guitar Tone Transformation Using GANs and Integrated Unaligned Clean Data
by: Chen, Yu-Hua, et al.
Published: (2024)
by: Chen, Yu-Hua, et al.
Published: (2024)
Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models
by: Simon, Christian, et al.
Published: (2026)
by: Simon, Christian, et al.
Published: (2026)
WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
by: Yang, Qihui, et al.
Published: (2025)
by: Yang, Qihui, et al.
Published: (2025)
VIRTUE: Visual-Interactive Text-Image Universal Embedder
by: Wang, Wei-Yao, et al.
Published: (2025)
by: Wang, Wei-Yao, et al.
Published: (2025)
Type II Lifshitz invariant and optically active Higgs mode in time-reversal symmetry broken superconductors
by: Nagashima, Raigo, et al.
Published: (2026)
by: Nagashima, Raigo, et al.
Published: (2026)
Visual Echoes: A Simple Unified Transformer for Audio-Visual Generation
by: Yang, Shiqi, et al.
Published: (2024)
by: Yang, Shiqi, et al.
Published: (2024)
SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
by: Shimada, Kazuki, et al.
Published: (2024)
by: Shimada, Kazuki, et al.
Published: (2024)
On the de-duplication of the Lakh MIDI dataset
by: Choi, Eunjin, et al.
Published: (2025)
by: Choi, Eunjin, et al.
Published: (2025)
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
by: Zhang, Yixiao, et al.
Published: (2024)
by: Zhang, Yixiao, et al.
Published: (2024)
Similar Items
-
Variable Bitrate Residual Vector Quantization for Audio Coding
by: Chae, Yunkee, et al.
Published: (2024) -
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
by: Choi, Woosung, et al.
Published: (2025) -
Noise-to-Notes: Diffusion-based Generation and Refinement for Automatic Drum Transcription
by: Yeung, Michael, et al.
Published: (2025) -
Do Foundational Audio Encoders Understand Music Structure?
by: Toyama, Keisuke, et al.
Published: (2025) -
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
by: Comunità, Marco, et al.
Published: (2024)