BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Qian, Lekai, Gu, Haoyu, Zhao, Jingwei, Wang, Ziyu
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918534863388672
author Qian, Lekai
Gu, Haoyu
Zhao, Jingwei
Wang, Ziyu
author_facet Qian, Lekai
Gu, Haoyu
Zhao, Jingwei
Wang, Ziyu
contents Tokenizing music to fit the general framework of language models is a compelling challenge, especially considering the diverse symbolic structures in which music can be represented (e.g., sequences, grids, and graphs). To date, most approaches tokenize symbolic music as sequences of musical events, such as onsets, pitches, time shifts, or compound note events. This strategy is intuitive and has proven effective in Transformer-based models, but it treats the regularity of musical time implicitly: individual tokens may span different durations, resulting in non-uniform time progression. In this paper, we instead consider whether an alternative tokenization is possible, where a uniform-length musical step (e.g., a beat) serves as the basic unit. Specifically, we encode all events within a single time step at the same pitch as one token, and group tokens explicitly by time step, which resembles a sparse encoding of a piano-roll representation. We evaluate the proposed tokenization on music continuation and accompaniment generation tasks, comparing it with mainstream event-based methods. Results show improved musical quality and structural coherence, while additional analyses confirm higher efficiency and more effective capture of long-range patterns with the proposed tokenization.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19532
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps
Qian, Lekai
Gu, Haoyu
Zhao, Jingwei
Wang, Ziyu
Sound
Artificial Intelligence
I.2.7; H.5.5
Tokenizing music to fit the general framework of language models is a compelling challenge, especially considering the diverse symbolic structures in which music can be represented (e.g., sequences, grids, and graphs). To date, most approaches tokenize symbolic music as sequences of musical events, such as onsets, pitches, time shifts, or compound note events. This strategy is intuitive and has proven effective in Transformer-based models, but it treats the regularity of musical time implicitly: individual tokens may span different durations, resulting in non-uniform time progression. In this paper, we instead consider whether an alternative tokenization is possible, where a uniform-length musical step (e.g., a beat) serves as the basic unit. Specifically, we encode all events within a single time step at the same pitch as one token, and group tokens explicitly by time step, which resembles a sparse encoding of a piano-roll representation. We evaluate the proposed tokenization on music continuation and accompaniment generation tasks, comparing it with mainstream event-based methods. Results show improved musical quality and structural coherence, while additional analyses confirm higher efficiency and more effective capture of long-range patterns with the proposed tokenization.
title BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps
topic Sound
Artificial Intelligence
I.2.7; H.5.5
url https://arxiv.org/abs/2604.19532