Annotation-Free MIDI-to-Audio Synthesis via Concatenative Synthesis and Generative Refinement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Take, Osamu, Akama, Taketo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912423312621568
author Take, Osamu
Akama, Taketo
author_facet Take, Osamu
Akama, Taketo
contents Recent MIDI-to-audio synthesis methods using deep neural networks have successfully generated high-quality, expressive instrumental tracks. However, these methods require MIDI annotations for supervised training, limiting the diversity of instrument timbres and expression styles in the output. We propose CoSaRef, a MIDI-to-audio synthesis method that does not require MIDI-audio paired datasets. CoSaRef first generates a synthetic audio track using concatenative synthesis based on MIDI input, then refines it with a diffusion-based deep generative model trained on datasets without MIDI annotations. This approach improves the diversity of timbres and expression styles. Additionally, it allows detailed control over timbres and expression through audio sample selection and extra MIDI design, similar to traditional functions in digital audio workstations. Experiments showed that CoSaRef could generate realistic tracks while preserving fine-grained timbre control via one-shot samples. Moreover, despite not being supervised on MIDI annotation, CoSaRef outperformed the state-of-the-art timbre-controllable method based on MIDI supervision in both objective and subjective evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16785
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Annotation-Free MIDI-to-Audio Synthesis via Concatenative Synthesis and Generative Refinement
Take, Osamu
Akama, Taketo
Sound
Machine Learning
Audio and Speech Processing
Recent MIDI-to-audio synthesis methods using deep neural networks have successfully generated high-quality, expressive instrumental tracks. However, these methods require MIDI annotations for supervised training, limiting the diversity of instrument timbres and expression styles in the output. We propose CoSaRef, a MIDI-to-audio synthesis method that does not require MIDI-audio paired datasets. CoSaRef first generates a synthetic audio track using concatenative synthesis based on MIDI input, then refines it with a diffusion-based deep generative model trained on datasets without MIDI annotations. This approach improves the diversity of timbres and expression styles. Additionally, it allows detailed control over timbres and expression through audio sample selection and extra MIDI design, similar to traditional functions in digital audio workstations. Experiments showed that CoSaRef could generate realistic tracks while preserving fine-grained timbre control via one-shot samples. Moreover, despite not being supervised on MIDI annotation, CoSaRef outperformed the state-of-the-art timbre-controllable method based on MIDI supervision in both objective and subjective evaluation.
title Annotation-Free MIDI-to-Audio Synthesis via Concatenative Synthesis and Generative Refinement
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2410.16785