Saved in:
Bibliographic Details
Main Authors: Neplenbroek, Vera, Bisazza, Arianna, Fernández, Raquel
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.16467
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909790967431168
author Neplenbroek, Vera
Bisazza, Arianna
Fernández, Raquel
author_facet Neplenbroek, Vera
Bisazza, Arianna
Fernández, Raquel
contents Generative Large Language Models (LLMs) infer user's demographic information from subtle cues in the conversation -- a phenomenon called implicit personalization. Prior work has shown that such inferences can lead to lower quality responses for users assumed to be from minority groups, even when no demographic information is explicitly provided. In this work, we systematically explore how LLMs respond to stereotypical cues using controlled synthetic conversations, by analyzing the models' latent user representations through both model internals and generated answers to targeted user questions. Our findings reveal that LLMs do infer demographic attributes based on these stereotypical signals, which for a number of groups even persists when the user explicitly identifies with a different demographic group. Finally, we show that this form of stereotype-driven implicit personalization can be effectively mitigated by intervening on the model's internal representations using a trained linear probe to steer them toward the explicitly stated identity. Our results highlight the need for greater transparency and control in how LLMs represent user identity.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16467
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization
Neplenbroek, Vera
Bisazza, Arianna
Fernández, Raquel
Computation and Language
Generative Large Language Models (LLMs) infer user's demographic information from subtle cues in the conversation -- a phenomenon called implicit personalization. Prior work has shown that such inferences can lead to lower quality responses for users assumed to be from minority groups, even when no demographic information is explicitly provided. In this work, we systematically explore how LLMs respond to stereotypical cues using controlled synthetic conversations, by analyzing the models' latent user representations through both model internals and generated answers to targeted user questions. Our findings reveal that LLMs do infer demographic attributes based on these stereotypical signals, which for a number of groups even persists when the user explicitly identifies with a different demographic group. Finally, we show that this form of stereotype-driven implicit personalization can be effectively mitigated by intervening on the model's internal representations using a trained linear probe to steer them toward the explicitly stated identity. Our results highlight the need for greater transparency and control in how LLMs represent user identity.
title Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization
topic Computation and Language
url https://arxiv.org/abs/2505.16467