ST-ITO: Controlling Audio Effects for Style Transfer with Inference-Time Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Steinmetz, Christian J., Singh, Shubhr, Comunità, Marco, Ibnyahya, Ilias, Yuan, Shanxin, Benetos, Emmanouil, Reiss, Joshua D.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916457122627584
author Steinmetz, Christian J.
Singh, Shubhr
Comunità, Marco
Ibnyahya, Ilias
Yuan, Shanxin
Benetos, Emmanouil
Reiss, Joshua D.
author_facet Steinmetz, Christian J.
Singh, Shubhr
Comunità, Marco
Ibnyahya, Ilias
Yuan, Shanxin
Benetos, Emmanouil
Reiss, Joshua D.
contents Audio production style transfer is the task of processing an input to impart stylistic elements from a reference recording. Existing approaches often train a neural network to estimate control parameters for a set of audio effects. However, these approaches are limited in that they can only control a fixed set of effects, where the effects must be differentiable or otherwise employ specialized training techniques. In this work, we introduce ST-ITO, Style Transfer with Inference-Time Optimization, an approach that instead searches the parameter space of an audio effect chain at inference. This method enables control of arbitrary audio effect chains, including unseen and non-differentiable effects. Our approach employs a learned metric of audio production style, which we train through a simple and scalable self-supervised pretraining strategy, along with a gradient-free optimizer. Due to the limited existing evaluation methods for audio production style transfer, we introduce a multi-part benchmark to evaluate audio production style metrics and style transfer systems. This evaluation demonstrates that our audio representation better captures attributes related to audio production and enables expressive style transfer via control of arbitrary audio effects.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21233
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ST-ITO: Controlling Audio Effects for Style Transfer with Inference-Time Optimization
Steinmetz, Christian J.
Singh, Shubhr
Comunità, Marco
Ibnyahya, Ilias
Yuan, Shanxin
Benetos, Emmanouil
Reiss, Joshua D.
Sound
Audio and Speech Processing
Audio production style transfer is the task of processing an input to impart stylistic elements from a reference recording. Existing approaches often train a neural network to estimate control parameters for a set of audio effects. However, these approaches are limited in that they can only control a fixed set of effects, where the effects must be differentiable or otherwise employ specialized training techniques. In this work, we introduce ST-ITO, Style Transfer with Inference-Time Optimization, an approach that instead searches the parameter space of an audio effect chain at inference. This method enables control of arbitrary audio effect chains, including unseen and non-differentiable effects. Our approach employs a learned metric of audio production style, which we train through a simple and scalable self-supervised pretraining strategy, along with a gradient-free optimizer. Due to the limited existing evaluation methods for audio production style transfer, we introduce a multi-part benchmark to evaluate audio production style metrics and style transfer systems. This evaluation demonstrates that our audio representation better captures attributes related to audio production and enables expressive style transfer via control of arbitrary audio effects.
title ST-ITO: Controlling Audio Effects for Style Transfer with Inference-Time Optimization
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.21233