The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: O'Reilly, Patrick, Barnett, Julia, García, Hugo Flores, Chu, Annie, Pruyne, Nathan, Seetharaman, Prem, Pardo, Bryan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916957562863616
author O'Reilly, Patrick
Barnett, Julia
García, Hugo Flores
Chu, Annie
Pruyne, Nathan
Seetharaman, Prem
Pardo, Bryan
author_facet O'Reilly, Patrick
Barnett, Julia
García, Hugo Flores
Chu, Annie
Pruyne, Nathan
Seetharaman, Prem
Pardo, Bryan
contents Musicians and nonmusicians alike use rhythmic sound gestures, such as tapping and beatboxing, to express drum patterns. While these gestures effectively communicate musical ideas, realizing these ideas as fully-produced drum recordings can be time-consuming, potentially disrupting many creative workflows. To bridge this gap, we present TRIA (The Rhythm In Anything), a masked transformer model for mapping rhythmic sound gestures to high-fidelity drum recordings. Given an audio prompt of the desired rhythmic pattern and a second prompt to represent drumkit timbre, TRIA produces audio of a drumkit playing the desired rhythm (with appropriate elaborations) in the desired timbre. Subjective and objective evaluations show that a TRIA model trained on less than 10 hours of publicly-available drum data can generate high-quality, faithful realizations of sound gestures across a wide range of timbres in a zero-shot manner.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15625
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
O'Reilly, Patrick
Barnett, Julia
García, Hugo Flores
Chu, Annie
Pruyne, Nathan
Seetharaman, Prem
Pardo, Bryan
Sound
Audio and Speech Processing
Musicians and nonmusicians alike use rhythmic sound gestures, such as tapping and beatboxing, to express drum patterns. While these gestures effectively communicate musical ideas, realizing these ideas as fully-produced drum recordings can be time-consuming, potentially disrupting many creative workflows. To bridge this gap, we present TRIA (The Rhythm In Anything), a masked transformer model for mapping rhythmic sound gestures to high-fidelity drum recordings. Given an audio prompt of the desired rhythmic pattern and a second prompt to represent drumkit timbre, TRIA produces audio of a drumkit playing the desired rhythm (with appropriate elaborations) in the desired timbre. Subjective and objective evaluations show that a TRIA model trained on less than 10 hours of publicly-available drum data can generate high-quality, faithful realizations of sound gestures across a wide range of timbres in a zero-shot manner.
title The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.15625