Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Xianwei, Quan, Dou, Zhang, Zhenliang, Wang, Shuang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918406208356352
author Cao, Xianwei
Quan, Dou
Zhang, Zhenliang
Wang, Shuang
author_facet Cao, Xianwei
Quan, Dou
Zhang, Zhenliang
Wang, Shuang
contents Humans often juggle multiple, sometimes conflicting objectives and shift their priorities as circumstances change, rather than following a fixed objective function. In contrast, most computational decision-making and multi-objective RL methods assume static preference weights or a known scalar reward. In this work, we study sequential decision-making problem when these preference weights are unobserved latent variables that drift with context. Specifically, we propose Dynamic Preference Inference (DPI), a cognitively inspired framework in which an agent maintains a probabilistic belief over preference weights, updates this belief from recent interaction, and conditions its policy on inferred preferences. We instantiate DPI as a variational preference inference module trained jointly with a preference-conditioned actor-critic, using vector-valued returns as evidence about latent trade-offs. In queueing, maze, and multi-objective continuous-control environments with event-driven changes in objectives, DPI adapts its inferred preferences to new regimes and achieves higher post-shift performance than fixed-weight and heuristic envelope baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2603_22813
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts
Cao, Xianwei
Quan, Dou
Zhang, Zhenliang
Wang, Shuang
Artificial Intelligence
Humans often juggle multiple, sometimes conflicting objectives and shift their priorities as circumstances change, rather than following a fixed objective function. In contrast, most computational decision-making and multi-objective RL methods assume static preference weights or a known scalar reward. In this work, we study sequential decision-making problem when these preference weights are unobserved latent variables that drift with context. Specifically, we propose Dynamic Preference Inference (DPI), a cognitively inspired framework in which an agent maintains a probabilistic belief over preference weights, updates this belief from recent interaction, and conditions its policy on inferred preferences. We instantiate DPI as a variational preference inference module trained jointly with a preference-conditioned actor-critic, using vector-valued returns as evidence about latent trade-offs. In queueing, maze, and multi-objective continuous-control environments with event-driven changes in objectives, DPI adapts its inferred preferences to new regimes and achieves higher post-shift performance than fixed-weight and heuristic envelope baselines.
title Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts
topic Artificial Intelligence
url https://arxiv.org/abs/2603.22813