Flexible Agent Alignment with Goal Inference from Open-Ended Dialog

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Rachel, Qu, Jingyi, Bobu, Andreea, Hadfield-Menell, Dylan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913097666527232
author Ma, Rachel
Qu, Jingyi
Bobu, Andreea
Hadfield-Menell, Dylan
author_facet Ma, Rachel
Qu, Jingyi
Bobu, Andreea
Hadfield-Menell, Dylan
contents We introduce Open-Universe Assistance Games (OU-AGs), a formal framework extending assistance games to LLM-based agents. Effective assistance requires reasoning over human preferences that are unbounded, underspecified, and evolving. Current LLM agents struggle in multi-turn interactions and with maintaining accurate models of user intent in collaborative settings. Existing assistance game formulations assume fixed, predefined preferences, an assumption that breaks down in open-ended dialogue where goals are revised incrementally and expressed in natural language. Grounded in cognitive science accounts of preference construction, we represent human preferences as a dynamically updated distribution over discrete natural-language goals. To operationalize OU-AGs, we introduce GOOD (GOals from Open-ended Dialogue), a data-efficient online method that extracts and ranks candidate goals during interaction, using LLM-simulated users to perform probabilistic inference over goal hypotheses. This allows for interpretable, uncertainty-aware preference representations without large offline datasets. We evaluate GOOD across three text-based domains: grocery shopping, household robotics (AI2-THOR), and coding. Compared to baselines without explicit goal tracking, GOOD produces semantically coherent goal representations and improves alignment with user intent across domains.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15119
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
Ma, Rachel
Qu, Jingyi
Bobu, Andreea
Hadfield-Menell, Dylan
Artificial Intelligence
Computation and Language
Machine Learning
Robotics
We introduce Open-Universe Assistance Games (OU-AGs), a formal framework extending assistance games to LLM-based agents. Effective assistance requires reasoning over human preferences that are unbounded, underspecified, and evolving. Current LLM agents struggle in multi-turn interactions and with maintaining accurate models of user intent in collaborative settings. Existing assistance game formulations assume fixed, predefined preferences, an assumption that breaks down in open-ended dialogue where goals are revised incrementally and expressed in natural language. Grounded in cognitive science accounts of preference construction, we represent human preferences as a dynamically updated distribution over discrete natural-language goals. To operationalize OU-AGs, we introduce GOOD (GOals from Open-ended Dialogue), a data-efficient online method that extracts and ranks candidate goals during interaction, using LLM-simulated users to perform probabilistic inference over goal hypotheses. This allows for interpretable, uncertainty-aware preference representations without large offline datasets. We evaluate GOOD across three text-based domains: grocery shopping, household robotics (AI2-THOR), and coding. Compared to baselines without explicit goal tracking, GOOD produces semantically coherent goal representations and improves alignment with user intent across domains.
title Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
topic Artificial Intelligence
Computation and Language
Machine Learning
Robotics
url https://arxiv.org/abs/2508.15119