GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Guinot, Julien, Quinton, Elio, Fazekas, György
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912445907337216
author Guinot, Julien
Quinton, Elio
Fazekas, György
author_facet Guinot, Julien
Quinton, Elio
Fazekas, György
contents Multimodal contrastive models have achieved strong performance in text-audio retrieval and zero-shot settings, but improving joint embedding spaces remains an active research area. Less attention has been given to making these systems controllable and interactive for users. In text-music retrieval, the ambiguity of freeform language creates a many-to-many mapping, often resulting in inflexible or unsatisfying results. We introduce Generative Diffusion Retriever (GDR), a novel framework that leverages diffusion models to generate queries in a retrieval-optimized latent space. This enables controllability through generative tools such as negative prompting and denoising diffusion implicit models (DDIM) inversion, opening a new direction in retrieval control. GDR improves retrieval performance over contrastive teacher models and supports retrieval in audio-only latent spaces using non-jointly trained encoders. Finally, we demonstrate that GDR enables effective post-hoc manipulation of retrieval behavior, enhancing interactive control for text-music retrieval tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17886
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models
Guinot, Julien
Quinton, Elio
Fazekas, György
Sound
Audio and Speech Processing
Multimodal contrastive models have achieved strong performance in text-audio retrieval and zero-shot settings, but improving joint embedding spaces remains an active research area. Less attention has been given to making these systems controllable and interactive for users. In text-music retrieval, the ambiguity of freeform language creates a many-to-many mapping, often resulting in inflexible or unsatisfying results. We introduce Generative Diffusion Retriever (GDR), a novel framework that leverages diffusion models to generate queries in a retrieval-optimized latent space. This enables controllability through generative tools such as negative prompting and denoising diffusion implicit models (DDIM) inversion, opening a new direction in retrieval control. GDR improves retrieval performance over contrastive teacher models and supports retrieval in audio-only latent spaces using non-jointly trained encoders. Finally, we demonstrate that GDR enables effective post-hoc manipulation of retrieval behavior, enhancing interactive control for text-music retrieval tasks.
title GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.17886