SynthCloner: Synthesizer-style Audio Transfer via Factorized Codec with ADSR Envelope Control
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911408038346752 |
|---|---|
| author | Liu, Jeng-Yue Hsu, Ting-Chao Yeh, Yen-Tung Su, Li Yang, Yi-Hsuan |
| author_facet | Liu, Jeng-Yue Hsu, Ting-Chao Yeh, Yen-Tung Su, Li Yang, Yi-Hsuan |
| contents | Electronic synthesizer sounds are controlled by parameter settings that yield complex timbral characteristics and ADSR envelopes, making synthesizer-style audio transfer particularly challenging. Recent approaches to timbre transfer often rely on spectral objectives or implicit style matching, offering limited control over envelope shaping. Moreover, public synthesizer datasets rarely provide diverse coverage of timbres and ADSR envelopes. To address these gaps, we present SynthCloner, a factorized codec model that disentangles audio into three attributes: ADSR envelope, timbre, and content. This separation enables expressive audio transfer with independent control over these attributes. Additionally, we introduce SynthCAT, a new synthesizer dataset with a task-specific rendering pipeline covering 250 timbres, 120 ADSR envelopes, and 100 MIDI sequences. Experiments show that SynthCloner outperforms baselines on both objective and subjective metrics, while enabling independent attribute control. The code, model checkpoint, and audio examples are available at https://buffett0323.github.io/synthcloner/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_24286 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | SynthCloner: Synthesizer-style Audio Transfer via Factorized Codec with ADSR Envelope Control Liu, Jeng-Yue Hsu, Ting-Chao Yeh, Yen-Tung Su, Li Yang, Yi-Hsuan Audio and Speech Processing Sound Electronic synthesizer sounds are controlled by parameter settings that yield complex timbral characteristics and ADSR envelopes, making synthesizer-style audio transfer particularly challenging. Recent approaches to timbre transfer often rely on spectral objectives or implicit style matching, offering limited control over envelope shaping. Moreover, public synthesizer datasets rarely provide diverse coverage of timbres and ADSR envelopes. To address these gaps, we present SynthCloner, a factorized codec model that disentangles audio into three attributes: ADSR envelope, timbre, and content. This separation enables expressive audio transfer with independent control over these attributes. Additionally, we introduce SynthCAT, a new synthesizer dataset with a task-specific rendering pipeline covering 250 timbres, 120 ADSR envelopes, and 100 MIDI sequences. Experiments show that SynthCloner outperforms baselines on both objective and subjective metrics, while enabling independent attribute control. The code, model checkpoint, and audio examples are available at https://buffett0323.github.io/synthcloner/. |
| title | SynthCloner: Synthesizer-style Audio Transfer via Factorized Codec with ADSR Envelope Control |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2509.24286 |