HorizonBench: Long-Horizon Personalization with Evolving Preferences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Shuyue Stella, Paranjape, Bhargavi, Oktar, Kerem, Ma, Zhongyao, Zhou, Gelin, Guan, Lin, Zhang, Na, Park, Sem, Chen, Lin, Yang, Diyi, Tsvetkov, Yulia, Celikyilmaz, Asli
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915944463335424
author Li, Shuyue Stella
Paranjape, Bhargavi
Oktar, Kerem
Ma, Zhongyao
Zhou, Gelin
Guan, Lin
Zhang, Na
Park, Sem
Chen, Lin
Yang, Diyi
Tsvetkov, Yulia
Celikyilmaz, Asli
author_facet Li, Shuyue Stella
Paranjape, Bhargavi
Oktar, Kerem
Ma, Zhongyao
Zhou, Gelin
Guan, Lin
Zhang, Na
Park, Sem
Chen, Lin
Yang, Diyi
Tsvetkov, Yulia
Celikyilmaz, Asli
contents User preferences evolve across months of interaction, and tracking them requires inferring when a stated preference has been changed by a subsequent life event. We define this problem as long-horizon personalization and observe that progress on it is limited by data availability and measurement, with no existing resource providing both naturalistic long-horizon interactions and the ground-truth provenance needed to diagnose why models fail. We introduce a data generator that produces conversations from a structured mental state graph, yielding ground-truth provenance for every preference change across 6-month timelines, and from it construct HorizonBench, a benchmark of 4,245 items from 360 simulated users with 6-month conversation histories averaging ~4,300 turns and ~163K tokens. HorizonBench provides a testbed for long-context modeling, memory-augmented architectures, theory-of-mind reasoning, and user modeling. Across 25 frontier models, the best model reaches 52.8% and most score at or below the 20% chance baseline. When these models err on evolved preferences, over a third of the time they select the user's originally stated value without tracking the updated user state. This belief-update failure persists across context lengths and expression explicitness levels, identifying state-tracking capability as the primary bottleneck for long-horizon personalization.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17283
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HorizonBench: Long-Horizon Personalization with Evolving Preferences
Li, Shuyue Stella
Paranjape, Bhargavi
Oktar, Kerem
Ma, Zhongyao
Zhou, Gelin
Guan, Lin
Zhang, Na
Park, Sem
Chen, Lin
Yang, Diyi
Tsvetkov, Yulia
Celikyilmaz, Asli
Computation and Language
Artificial Intelligence
User preferences evolve across months of interaction, and tracking them requires inferring when a stated preference has been changed by a subsequent life event. We define this problem as long-horizon personalization and observe that progress on it is limited by data availability and measurement, with no existing resource providing both naturalistic long-horizon interactions and the ground-truth provenance needed to diagnose why models fail. We introduce a data generator that produces conversations from a structured mental state graph, yielding ground-truth provenance for every preference change across 6-month timelines, and from it construct HorizonBench, a benchmark of 4,245 items from 360 simulated users with 6-month conversation histories averaging ~4,300 turns and ~163K tokens. HorizonBench provides a testbed for long-context modeling, memory-augmented architectures, theory-of-mind reasoning, and user modeling. Across 25 frontier models, the best model reaches 52.8% and most score at or below the 20% chance baseline. When these models err on evolved preferences, over a third of the time they select the user's originally stated value without tracking the updated user state. This belief-update failure persists across context lengths and expression explicitness levels, identifying state-tracking capability as the primary bottleneck for long-horizon personalization.
title HorizonBench: Long-Horizon Personalization with Evolving Preferences
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.17283