SynthCloner: Synthesizer-style Audio Transfer via Factorized Codec with ADSR Envelope Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jeng-Yue, Hsu, Ting-Chao, Yeh, Yen-Tung, Su, Li, Yang, Yi-Hsuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911408038346752
author Liu, Jeng-Yue
Hsu, Ting-Chao
Yeh, Yen-Tung
Su, Li
Yang, Yi-Hsuan
author_facet Liu, Jeng-Yue
Hsu, Ting-Chao
Yeh, Yen-Tung
Su, Li
Yang, Yi-Hsuan
contents Electronic synthesizer sounds are controlled by parameter settings that yield complex timbral characteristics and ADSR envelopes, making synthesizer-style audio transfer particularly challenging. Recent approaches to timbre transfer often rely on spectral objectives or implicit style matching, offering limited control over envelope shaping. Moreover, public synthesizer datasets rarely provide diverse coverage of timbres and ADSR envelopes. To address these gaps, we present SynthCloner, a factorized codec model that disentangles audio into three attributes: ADSR envelope, timbre, and content. This separation enables expressive audio transfer with independent control over these attributes. Additionally, we introduce SynthCAT, a new synthesizer dataset with a task-specific rendering pipeline covering 250 timbres, 120 ADSR envelopes, and 100 MIDI sequences. Experiments show that SynthCloner outperforms baselines on both objective and subjective metrics, while enabling independent attribute control. The code, model checkpoint, and audio examples are available at https://buffett0323.github.io/synthcloner/.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24286
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SynthCloner: Synthesizer-style Audio Transfer via Factorized Codec with ADSR Envelope Control
Liu, Jeng-Yue
Hsu, Ting-Chao
Yeh, Yen-Tung
Su, Li
Yang, Yi-Hsuan
Audio and Speech Processing
Sound
Electronic synthesizer sounds are controlled by parameter settings that yield complex timbral characteristics and ADSR envelopes, making synthesizer-style audio transfer particularly challenging. Recent approaches to timbre transfer often rely on spectral objectives or implicit style matching, offering limited control over envelope shaping. Moreover, public synthesizer datasets rarely provide diverse coverage of timbres and ADSR envelopes. To address these gaps, we present SynthCloner, a factorized codec model that disentangles audio into three attributes: ADSR envelope, timbre, and content. This separation enables expressive audio transfer with independent control over these attributes. Additionally, we introduce SynthCAT, a new synthesizer dataset with a task-specific rendering pipeline covering 250 timbres, 120 ADSR envelopes, and 100 MIDI sequences. Experiments show that SynthCloner outperforms baselines on both objective and subjective metrics, while enabling independent attribute control. The code, model checkpoint, and audio examples are available at https://buffett0323.github.io/synthcloner/.
title SynthCloner: Synthesizer-style Audio Transfer via Factorized Codec with ADSR Envelope Control
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2509.24286