TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Kyungsu, Koo, Junghyun, Lee, Sungho, Joung, Haesun, Lee, Kyogu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915149487538176
author Kim, Kyungsu
Koo, Junghyun
Lee, Sungho
Joung, Haesun
Lee, Kyogu
author_facet Kim, Kyungsu
Koo, Junghyun
Lee, Sungho
Joung, Haesun
Lee, Kyogu
contents Recent advancements in neural audio codecs have enabled the use of tokenized audio representations in various audio generation tasks, such as text-to-speech, text-to-audio, and text-to-music generation. Leveraging this approach, we propose TokenSynth, a novel neural synthesizer that utilizes a decoder-only transformer to generate desired audio tokens from MIDI tokens and CLAP (Contrastive Language-Audio Pretraining) embedding, which has timbre-related information. Our model is capable of performing instrument cloning, text-to-instrument synthesis, and text-guided timbre manipulation without any fine-tuning. This flexibility enables diverse sound design and intuitive timbre control. We evaluated the quality of the synthesized audio, the timbral similarity between synthesized and target audio/text, and synthesis accuracy (i.e., how accurately it follows the input MIDI) using objective measures. TokenSynth demonstrates the potential of leveraging advanced neural audio codecs and transformers to create powerful and versatile neural synthesizers. The source code, model weights, and audio demos are available at: https://github.com/KyungsuKim42/tokensynth
format Preprint
id arxiv_https___arxiv_org_abs_2502_08939
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument
Kim, Kyungsu
Koo, Junghyun
Lee, Sungho
Joung, Haesun
Lee, Kyogu
Sound
Artificial Intelligence
Recent advancements in neural audio codecs have enabled the use of tokenized audio representations in various audio generation tasks, such as text-to-speech, text-to-audio, and text-to-music generation. Leveraging this approach, we propose TokenSynth, a novel neural synthesizer that utilizes a decoder-only transformer to generate desired audio tokens from MIDI tokens and CLAP (Contrastive Language-Audio Pretraining) embedding, which has timbre-related information. Our model is capable of performing instrument cloning, text-to-instrument synthesis, and text-guided timbre manipulation without any fine-tuning. This flexibility enables diverse sound design and intuitive timbre control. We evaluated the quality of the synthesized audio, the timbral similarity between synthesized and target audio/text, and synthesis accuracy (i.e., how accurately it follows the input MIDI) using objective measures. TokenSynth demonstrates the potential of leveraging advanced neural audio codecs and transformers to create powerful and versatile neural synthesizers. The source code, model weights, and audio demos are available at: https://github.com/KyungsuKim42/tokensynth
title TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2502.08939