Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tsaknakis, Ioannis, Song, Bingqing, Gan, Shuyu, Kang, Dongyeop, Garcia, Alfredo, Liu, Gaowen, Fleming, Charles, Hong, Mingyi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911220563443712
author Tsaknakis, Ioannis
Song, Bingqing
Gan, Shuyu
Kang, Dongyeop
Garcia, Alfredo
Liu, Gaowen
Fleming, Charles
Hong, Mingyi
author_facet Tsaknakis, Ioannis
Song, Bingqing
Gan, Shuyu
Kang, Dongyeop
Garcia, Alfredo
Liu, Gaowen
Fleming, Charles
Hong, Mingyi
contents Large Language Models (LLMs) excel at producing broadly relevant text, but this generality becomes a limitation when user-specific preferences are required, such as recommending restaurants or planning travel. In these scenarios, users rarely articulate every preference explicitly; instead, much of what they care about remains latent, waiting to be inferred. This raises a fundamental question: Can LLMs uncover and reason about such latent information through conversation? We address this problem by introducing a unified benchmark for evaluating latent information discovery - the ability of LLMs to reveal and utilize hidden user attributes through multi-turn interaction. The benchmark spans three progressively realistic settings: the classic 20 Questions game, Personalized Question Answering, and Personalized Text Summarization. All tasks share a tri-agent framework (User, Assistant, Judge) enabling turn-level evaluation of elicitation and adaptation. Our results reveal that while LLMs can indeed surface latent information through dialogue, their success varies dramatically with context: from 32% to 98%, depending on task complexity, topic, and number of hidden attributes. This benchmark provides the first systematic framework for studying latent information discovery in personalized interaction, highlighting that effective preference inference remains an open frontier for building truly adaptive AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17132
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
Tsaknakis, Ioannis
Song, Bingqing
Gan, Shuyu
Kang, Dongyeop
Garcia, Alfredo
Liu, Gaowen
Fleming, Charles
Hong, Mingyi
Machine Learning
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) excel at producing broadly relevant text, but this generality becomes a limitation when user-specific preferences are required, such as recommending restaurants or planning travel. In these scenarios, users rarely articulate every preference explicitly; instead, much of what they care about remains latent, waiting to be inferred. This raises a fundamental question: Can LLMs uncover and reason about such latent information through conversation? We address this problem by introducing a unified benchmark for evaluating latent information discovery - the ability of LLMs to reveal and utilize hidden user attributes through multi-turn interaction. The benchmark spans three progressively realistic settings: the classic 20 Questions game, Personalized Question Answering, and Personalized Text Summarization. All tasks share a tri-agent framework (User, Assistant, Judge) enabling turn-level evaluation of elicitation and adaptation. Our results reveal that while LLMs can indeed surface latent information through dialogue, their success varies dramatically with context: from 32% to 98%, depending on task complexity, topic, and number of hidden attributes. This benchmark provides the first systematic framework for studying latent information discovery in personalized interaction, highlighting that effective preference inference remains an open frontier for building truly adaptive AI systems.
title Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.17132