Pragmatic Instruction Following and Goal Assistance via Cooperative Language-Guided Inverse Planning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhi-Xuan, Tan, Ying, Lance, Mansinghka, Vikash, Tenenbaum, Joshua B.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929258234904576
author Zhi-Xuan, Tan
Ying, Lance
Mansinghka, Vikash
Tenenbaum, Joshua B.
author_facet Zhi-Xuan, Tan
Ying, Lance
Mansinghka, Vikash
Tenenbaum, Joshua B.
contents People often give instructions whose meaning is ambiguous without further context, expecting that their actions or goals will disambiguate their intentions. How can we build assistive agents that follow such instructions in a flexible, context-sensitive manner? This paper introduces cooperative language-guided inverse plan search (CLIPS), a Bayesian agent architecture for pragmatic instruction following and goal assistance. Our agent assists a human by modeling them as a cooperative planner who communicates joint plans to the assistant, then performs multimodal Bayesian inference over the human's goal from actions and language, using large language models (LLMs) to evaluate the likelihood of an instruction given a hypothesized plan. Given this posterior, our assistant acts to minimize expected goal achievement cost, enabling it to pragmatically follow ambiguous instructions and provide effective assistance even when uncertain about the goal. We evaluate these capabilities in two cooperative planning domains (Doors, Keys & Gems and VirtualHome), finding that CLIPS significantly outperforms GPT-4V, LLM-based literal instruction following and unimodal inverse planning in both accuracy and helpfulness, while closely matching the inferences and assistive judgments provided by human raters.
format Preprint
id arxiv_https___arxiv_org_abs_2402_17930
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pragmatic Instruction Following and Goal Assistance via Cooperative Language-Guided Inverse Planning
Zhi-Xuan, Tan
Ying, Lance
Mansinghka, Vikash
Tenenbaum, Joshua B.
Artificial Intelligence
Computation and Language
Machine Learning
People often give instructions whose meaning is ambiguous without further context, expecting that their actions or goals will disambiguate their intentions. How can we build assistive agents that follow such instructions in a flexible, context-sensitive manner? This paper introduces cooperative language-guided inverse plan search (CLIPS), a Bayesian agent architecture for pragmatic instruction following and goal assistance. Our agent assists a human by modeling them as a cooperative planner who communicates joint plans to the assistant, then performs multimodal Bayesian inference over the human's goal from actions and language, using large language models (LLMs) to evaluate the likelihood of an instruction given a hypothesized plan. Given this posterior, our assistant acts to minimize expected goal achievement cost, enabling it to pragmatically follow ambiguous instructions and provide effective assistance even when uncertain about the goal. We evaluate these capabilities in two cooperative planning domains (Doors, Keys & Gems and VirtualHome), finding that CLIPS significantly outperforms GPT-4V, LLM-based literal instruction following and unimodal inverse planning in both accuracy and helpfulness, while closely matching the inferences and assistive judgments provided by human raters.
title Pragmatic Instruction Following and Goal Assistance via Cooperative Language-Guided Inverse Planning
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2402.17930