The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916957562863616 |
|---|---|
| author | O'Reilly, Patrick Barnett, Julia García, Hugo Flores Chu, Annie Pruyne, Nathan Seetharaman, Prem Pardo, Bryan |
| author_facet | O'Reilly, Patrick Barnett, Julia García, Hugo Flores Chu, Annie Pruyne, Nathan Seetharaman, Prem Pardo, Bryan |
| contents | Musicians and nonmusicians alike use rhythmic sound gestures, such as tapping and beatboxing, to express drum patterns. While these gestures effectively communicate musical ideas, realizing these ideas as fully-produced drum recordings can be time-consuming, potentially disrupting many creative workflows. To bridge this gap, we present TRIA (The Rhythm In Anything), a masked transformer model for mapping rhythmic sound gestures to high-fidelity drum recordings. Given an audio prompt of the desired rhythmic pattern and a second prompt to represent drumkit timbre, TRIA produces audio of a drumkit playing the desired rhythm (with appropriate elaborations) in the desired timbre. Subjective and objective evaluations show that a TRIA model trained on less than 10 hours of publicly-available drum data can generate high-quality, faithful realizations of sound gestures across a wide range of timbres in a zero-shot manner. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_15625 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling O'Reilly, Patrick Barnett, Julia García, Hugo Flores Chu, Annie Pruyne, Nathan Seetharaman, Prem Pardo, Bryan Sound Audio and Speech Processing Musicians and nonmusicians alike use rhythmic sound gestures, such as tapping and beatboxing, to express drum patterns. While these gestures effectively communicate musical ideas, realizing these ideas as fully-produced drum recordings can be time-consuming, potentially disrupting many creative workflows. To bridge this gap, we present TRIA (The Rhythm In Anything), a masked transformer model for mapping rhythmic sound gestures to high-fidelity drum recordings. Given an audio prompt of the desired rhythmic pattern and a second prompt to represent drumkit timbre, TRIA produces audio of a drumkit playing the desired rhythm (with appropriate elaborations) in the desired timbre. Subjective and objective evaluations show that a TRIA model trained on less than 10 hours of publicly-available drum data can generate high-quality, faithful realizations of sound gestures across a wide range of timbres in a zero-shot manner. |
| title | The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2509.15625 |