Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fabbri, Francesco, Penha, Gustavo, D'Amico, Edoardo, Wang, Alice, De Nadai, Marco, Doremus, Jackie, Gigioli, Paul, Damianou, Andreas, Stal, Oskar, Lalmas, Mounia
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915441604034560
author Fabbri, Francesco
Penha, Gustavo
D'Amico, Edoardo
Wang, Alice
De Nadai, Marco
Doremus, Jackie
Gigioli, Paul
Damianou, Andreas
Stal, Oskar
Lalmas, Mounia
author_facet Fabbri, Francesco
Penha, Gustavo
D'Amico, Edoardo
Wang, Alice
De Nadai, Marco
Doremus, Jackie
Gigioli, Paul
Damianou, Andreas
Stal, Oskar
Lalmas, Mounia
contents Evaluating personalized recommendations remains a central challenge, especially in long-form audio domains like podcasts, where traditional offline metrics suffer from exposure bias and online methods such as A/B testing are costly and operationally constrained. In this paper, we propose a novel framework that leverages Large Language Models (LLMs) as offline judges to assess the quality of podcast recommendations in a scalable and interpretable manner. Our two-stage profile-aware approach first constructs natural-language user profiles distilled from 90 days of listening history. These profiles summarize both topical interests and behavioral patterns, serving as compact, interpretable representations of user preferences. Rather than prompting the LLM with raw data, we use these profiles to provide high-level, semantically rich context-enabling the LLM to reason more effectively about alignment between a user's interests and recommended episodes. This reduces input complexity and improves interpretability. The LLM is then prompted to deliver fine-grained pointwise and pairwise judgments based on the profile-episode match. In a controlled study with 47 participants, our profile-aware judge matched human judgments with high fidelity and outperformed or matched a variant using raw listening histories. The framework enables efficient, profile-aware evaluation for iterative testing and model selection in recommender systems.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08777
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
Fabbri, Francesco
Penha, Gustavo
D'Amico, Edoardo
Wang, Alice
De Nadai, Marco
Doremus, Jackie
Gigioli, Paul
Damianou, Andreas
Stal, Oskar
Lalmas, Mounia
Information Retrieval
Artificial Intelligence
Machine Learning
Evaluating personalized recommendations remains a central challenge, especially in long-form audio domains like podcasts, where traditional offline metrics suffer from exposure bias and online methods such as A/B testing are costly and operationally constrained. In this paper, we propose a novel framework that leverages Large Language Models (LLMs) as offline judges to assess the quality of podcast recommendations in a scalable and interpretable manner. Our two-stage profile-aware approach first constructs natural-language user profiles distilled from 90 days of listening history. These profiles summarize both topical interests and behavioral patterns, serving as compact, interpretable representations of user preferences. Rather than prompting the LLM with raw data, we use these profiles to provide high-level, semantically rich context-enabling the LLM to reason more effectively about alignment between a user's interests and recommended episodes. This reduces input complexity and improves interpretability. The LLM is then prompted to deliver fine-grained pointwise and pairwise judgments based on the profile-episode match. In a controlled study with 47 participants, our profile-aware judge matched human judgments with high fidelity and outperformed or matched a variant using raw listening histories. The framework enables efficient, profile-aware evaluation for iterative testing and model selection in recommender systems.
title Evaluating Podcast Recommendations with Profile-Aware LLM-as-a-Judge
topic Information Retrieval
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.08777