MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Jingyue, Novack, Zachary, Long, Phillip, Hou, Yupeng, Chen, Ke, Berg-Kirkpatrick, Taylor, McAuley, Julian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909854312955904
author Huang, Jingyue
Novack, Zachary
Long, Phillip
Hou, Yupeng
Chen, Ke
Berg-Kirkpatrick, Taylor
McAuley, Julian
author_facet Huang, Jingyue
Novack, Zachary
Long, Phillip
Hou, Yupeng
Chen, Ke
Berg-Kirkpatrick, Taylor
McAuley, Julian
contents Discrete representation learning has shown promising results across various domains, including generation and understanding in image, speech and language. Inspired by these advances, we propose MuseTok, a tokenization method for symbolic music, and investigate its effectiveness in both music generation and understanding tasks. MuseTok employs the residual vector quantized-variational autoencoder (RQ-VAE) on bar-wise music segments within a Transformer-based encoder-decoder framework, producing music codes that achieve high-fidelity music reconstruction and accurate understanding of music theory. For comprehensive evaluation, we apply MuseTok to music generation and semantic understanding tasks, including melody extraction, chord recognition, and emotion recognition. Models incorporating MuseTok outperform previous representation learning baselines in semantic understanding while maintaining comparable performance in content generation. Furthermore, qualitative analyses on MuseTok codes, using ground-truth categories and synthetic datasets, reveal that MuseTok effectively captures underlying musical concepts from large music collections.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16273
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
Huang, Jingyue
Novack, Zachary
Long, Phillip
Hou, Yupeng
Chen, Ke
Berg-Kirkpatrick, Taylor
McAuley, Julian
Sound
Artificial Intelligence
Audio and Speech Processing
Discrete representation learning has shown promising results across various domains, including generation and understanding in image, speech and language. Inspired by these advances, we propose MuseTok, a tokenization method for symbolic music, and investigate its effectiveness in both music generation and understanding tasks. MuseTok employs the residual vector quantized-variational autoencoder (RQ-VAE) on bar-wise music segments within a Transformer-based encoder-decoder framework, producing music codes that achieve high-fidelity music reconstruction and accurate understanding of music theory. For comprehensive evaluation, we apply MuseTok to music generation and semantic understanding tasks, including melody extraction, chord recognition, and emotion recognition. Models incorporating MuseTok outperform previous representation learning baselines in semantic understanding while maintaining comparable performance in content generation. Furthermore, qualitative analyses on MuseTok codes, using ground-truth categories and synthetic datasets, reveal that MuseTok effectively captures underlying musical concepts from large music collections.
title MuseTok: Symbolic Music Tokenization for Generation and Semantic Understanding
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2510.16273