Benchmarking and Improving LLM Robustness for Personalized Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Okite, Chimaobi, Deng, Naihao, Bodipati, Kiran, Hou, Huaidian, Chai, Joyce, Mihalcea, Rada
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916965214322688
author Okite, Chimaobi
Deng, Naihao
Bodipati, Kiran
Hou, Huaidian
Chai, Joyce
Mihalcea, Rada
author_facet Okite, Chimaobi
Deng, Naihao
Bodipati, Kiran
Hou, Huaidian
Chai, Joyce
Mihalcea, Rada
contents Recent years have witnessed a growing interest in personalizing the responses of large language models (LLMs). While existing evaluations primarily focus on whether a response aligns with a user's preferences, we argue that factuality is an equally important yet often overlooked dimension. In the context of personalization, we define a model as robust if its responses are both factually accurate and align with the user preferences. To assess this, we introduce PERG, a scalable framework for evaluating robustness in LLMs, along with a new dataset, PERGData. We evaluate fourteen models from five different model families using different prompting methods. Our findings show that current LLMs struggle with robust personalization: even the strongest models (GPT-4.1, LLaMA3-70B) fail to maintain correctness in 5% of previously successful cases without personalization, while smaller models (e.g., 7B-scale) can fail more than 20% of the time. Further analysis reveals that robustness is significantly affected by the nature of the query and the type of user preference. To mitigate these failures, we propose Pref-Aligner, a two-stage approach that improves robustness by an average of 25% across models. Our work highlights critical gaps in current evaluation practices and introduces tools and metrics to support more reliable, user-aligned LLM deployments.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19358
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking and Improving LLM Robustness for Personalized Generation
Okite, Chimaobi
Deng, Naihao
Bodipati, Kiran
Hou, Huaidian
Chai, Joyce
Mihalcea, Rada
Computation and Language
Artificial Intelligence
Recent years have witnessed a growing interest in personalizing the responses of large language models (LLMs). While existing evaluations primarily focus on whether a response aligns with a user's preferences, we argue that factuality is an equally important yet often overlooked dimension. In the context of personalization, we define a model as robust if its responses are both factually accurate and align with the user preferences. To assess this, we introduce PERG, a scalable framework for evaluating robustness in LLMs, along with a new dataset, PERGData. We evaluate fourteen models from five different model families using different prompting methods. Our findings show that current LLMs struggle with robust personalization: even the strongest models (GPT-4.1, LLaMA3-70B) fail to maintain correctness in 5% of previously successful cases without personalization, while smaller models (e.g., 7B-scale) can fail more than 20% of the time. Further analysis reveals that robustness is significantly affected by the nature of the query and the type of user preference. To mitigate these failures, we propose Pref-Aligner, a two-stage approach that improves robustness by an average of 25% across models. Our work highlights critical gaps in current evaluation practices and introduces tools and metrics to support more reliable, user-aligned LLM deployments.
title Benchmarking and Improving LLM Robustness for Personalized Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.19358