Revealing Potential Biases in LLM-Based Recommender Systems in the Cold Start Setting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Andre, Alexandre, Roy, Gauthier, Dyer, Eva, Wang, Kai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916936889139200
author Andre, Alexandre
Roy, Gauthier
Dyer, Eva
Wang, Kai
author_facet Andre, Alexandre
Roy, Gauthier
Dyer, Eva
Wang, Kai
contents Large Language Models (LLMs) are increasingly used for recommendation tasks due to their general-purpose capabilities. While LLMs perform well in rich-context settings, their behavior in cold-start scenarios, where only limited signals such as age, gender, or language are available, raises fairness concerns because they may rely on societal biases encoded during pretraining. We introduce a benchmark specifically designed to evaluate fairness in zero-context recommendation. Our modular pipeline supports configurable recommendation domains and sensitive attributes, enabling systematic and flexible audits of any open-source LLM. Through evaluations of state-of-the-art models (Gemma 3 and Llama 3.2), we uncover consistent biases across recommendation domains (music, movies, and colleges) including gendered and cultural stereotypes. We also reveal a non-linear relationship between model size and fairness, highlighting the need for nuanced analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20401
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revealing Potential Biases in LLM-Based Recommender Systems in the Cold Start Setting
Andre, Alexandre
Roy, Gauthier
Dyer, Eva
Wang, Kai
Information Retrieval
Machine Learning
Large Language Models (LLMs) are increasingly used for recommendation tasks due to their general-purpose capabilities. While LLMs perform well in rich-context settings, their behavior in cold-start scenarios, where only limited signals such as age, gender, or language are available, raises fairness concerns because they may rely on societal biases encoded during pretraining. We introduce a benchmark specifically designed to evaluate fairness in zero-context recommendation. Our modular pipeline supports configurable recommendation domains and sensitive attributes, enabling systematic and flexible audits of any open-source LLM. Through evaluations of state-of-the-art models (Gemma 3 and Llama 3.2), we uncover consistent biases across recommendation domains (music, movies, and colleges) including gendered and cultural stereotypes. We also reveal a non-linear relationship between model size and fairness, highlighting the need for nuanced analysis.
title Revealing Potential Biases in LLM-Based Recommender Systems in the Cold Start Setting
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2508.20401