FxSearcher: gradient-free text-driven audio transformation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ki, Hojoon, Kim, Jongsuk, Kwon, Minchan, Kim, Junmo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912720231596032
author Ki, Hojoon
Kim, Jongsuk
Kwon, Minchan
Kim, Junmo
author_facet Ki, Hojoon
Kim, Jongsuk
Kwon, Minchan
Kim, Junmo
contents Achieving diverse and high-quality audio transformations from text prompts remains challenging, as existing methods are fundamentally constrained by their reliance on a limited set of differentiable audio effects. This paper proposes FxSearcher, a novel gradient-free framework that discovers the optimal configuration of audio effects (FX) to transform a source signal according to a text prompt. Our method employs Bayesian Optimization and CLAP-based score function to perform this search efficiently. Furthermore, a guiding prompt is introduced to prevent undesirable artifacts and enhance human preference. To objectively evaluate our method, we propose an AI-based evaluation framework. The results demonstrate that the highest scores achieved by our method on these metrics align closely with human preferences. Demos are available at https://hojoonki.github.io/FxSearcher/
format Preprint
id arxiv_https___arxiv_org_abs_2511_14138
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FxSearcher: gradient-free text-driven audio transformation
Ki, Hojoon
Kim, Jongsuk
Kwon, Minchan
Kim, Junmo
Audio and Speech Processing
Sound
Achieving diverse and high-quality audio transformations from text prompts remains challenging, as existing methods are fundamentally constrained by their reliance on a limited set of differentiable audio effects. This paper proposes FxSearcher, a novel gradient-free framework that discovers the optimal configuration of audio effects (FX) to transform a source signal according to a text prompt. Our method employs Bayesian Optimization and CLAP-based score function to perform this search efficiently. Furthermore, a guiding prompt is introduced to prevent undesirable artifacts and enhance human preference. To objectively evaluate our method, we propose an AI-based evaluation framework. The results demonstrate that the highest scores achieved by our method on these metrics align closely with human preferences. Demos are available at https://hojoonki.github.io/FxSearcher/
title FxSearcher: gradient-free text-driven audio transformation
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2511.14138