Text2Tracks: Prompt-based Music Recommendation via Generative Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Palumbo, Enrico, Penha, Gustavo, Damianou, Andreas, García, José Luis Redondo, Heath, Timothy Christopher, Wang, Alice, Bouchard, Hugues, Lalmas, Mounia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913772782747648
author Palumbo, Enrico
Penha, Gustavo
Damianou, Andreas
García, José Luis Redondo
Heath, Timothy Christopher
Wang, Alice
Bouchard, Hugues
Lalmas, Mounia
author_facet Palumbo, Enrico
Penha, Gustavo
Damianou, Andreas
García, José Luis Redondo
Heath, Timothy Christopher
Wang, Alice
Bouchard, Hugues
Lalmas, Mounia
contents In recent years, Large Language Models (LLMs) have enabled users to provide highly specific music recommendation requests using natural language prompts (e.g. "Can you recommend some old classics for slow dancing?"). In this setup, the recommended tracks are predicted by the LLM in an autoregressive way, i.e. the LLM generates the track titles one token at a time. While intuitive, this approach has several limitation. First, it is based on a general purpose tokenization that is optimized for words rather than for track titles. Second, it necessitates an additional entity resolution layer that matches the track title to the actual track identifier. Third, the number of decoding steps scales linearly with the length of the track title, slowing down inference. In this paper, we propose to address the task of prompt-based music recommendation as a generative retrieval task. Within this setting, we introduce novel, effective, and efficient representations of track identifiers that significantly outperform commonly used strategies. We introduce Text2Tracks, a generative retrieval model that learns a mapping from a user's music recommendation prompt to the relevant track IDs directly. Through an offline evaluation on a dataset of playlists with language inputs, we find that (1) the strategy to create IDs for music tracks is the most important factor for the effectiveness of Text2Tracks and semantic IDs significantly outperform commonly used strategies that rely on song titles as identifiers (2) provided with the right choice of track identifiers, Text2Tracks outperforms sparse and dense retrieval solutions trained to retrieve tracks from language prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2503_24193
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text2Tracks: Prompt-based Music Recommendation via Generative Retrieval
Palumbo, Enrico
Penha, Gustavo
Damianou, Andreas
García, José Luis Redondo
Heath, Timothy Christopher
Wang, Alice
Bouchard, Hugues
Lalmas, Mounia
Information Retrieval
In recent years, Large Language Models (LLMs) have enabled users to provide highly specific music recommendation requests using natural language prompts (e.g. "Can you recommend some old classics for slow dancing?"). In this setup, the recommended tracks are predicted by the LLM in an autoregressive way, i.e. the LLM generates the track titles one token at a time. While intuitive, this approach has several limitation. First, it is based on a general purpose tokenization that is optimized for words rather than for track titles. Second, it necessitates an additional entity resolution layer that matches the track title to the actual track identifier. Third, the number of decoding steps scales linearly with the length of the track title, slowing down inference. In this paper, we propose to address the task of prompt-based music recommendation as a generative retrieval task. Within this setting, we introduce novel, effective, and efficient representations of track identifiers that significantly outperform commonly used strategies. We introduce Text2Tracks, a generative retrieval model that learns a mapping from a user's music recommendation prompt to the relevant track IDs directly. Through an offline evaluation on a dataset of playlists with language inputs, we find that (1) the strategy to create IDs for music tracks is the most important factor for the effectiveness of Text2Tracks and semantic IDs significantly outperform commonly used strategies that rely on song titles as identifiers (2) provided with the right choice of track identifiers, Text2Tracks outperforms sparse and dense retrieval solutions trained to retrieve tracks from language prompts.
title Text2Tracks: Prompt-based Music Recommendation via Generative Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2503.24193