The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Angelina, Ho, Daniel E., Koyejo, Sanmi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911172281761792
author Wang, Angelina
Ho, Daniel E.
Koyejo, Sanmi
author_facet Wang, Angelina
Ho, Daniel E.
Koyejo, Sanmi
contents Standard offline evaluations for language models -- a series of independent, state-less inferences made by models -- fail to capture how language models actually behave in practice, where personalization fundamentally alters model behavior. For instance, identical benchmark questions to the same language model can produce markedly different responses when prompted to a state-less system, in one user's chat session, or in a different user's chat session. In this work, we provide empirical evidence showcasing this phenomenon by comparing offline evaluations to field evaluations conducted by having 800 real users of ChatGPT and Gemini pose benchmark and other provided questions to their chat interfaces.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19364
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
Wang, Angelina
Ho, Daniel E.
Koyejo, Sanmi
Computation and Language
Artificial Intelligence
Standard offline evaluations for language models -- a series of independent, state-less inferences made by models -- fail to capture how language models actually behave in practice, where personalization fundamentally alters model behavior. For instance, identical benchmark questions to the same language model can produce markedly different responses when prompted to a state-less system, in one user's chat session, or in a different user's chat session. In this work, we provide empirical evidence showcasing this phenomenon by comparing offline evaluations to field evaluations conducted by having 800 real users of ChatGPT and Gemini pose benchmark and other provided questions to their chat interfaces.
title The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.19364