PREDICT: Preference Reasoning by Evaluating Decomposed preferences Inferred from Candidate Trajectories
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917798012256256 |
|---|---|
| author | Aroca-Ouellette, Stephane Mackraz, Natalie Theobald, Barry-John Metcalf, Katherine |
| author_facet | Aroca-Ouellette, Stephane Mackraz, Natalie Theobald, Barry-John Metcalf, Katherine |
| contents | Accommodating human preferences is essential for creating AI agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs to infer preferences from user interactions, but they often produce broad and generic preferences, failing to capture the unique and individualized nature of human preferences. This paper introduces PREDICT, a method designed to enhance the precision and adaptability of inferring preferences. PREDICT incorporates three key elements: (1) iterative refinement of inferred preferences, (2) decomposition of preferences into constituent components, and (3) validation of preferences across multiple trajectories. We evaluate PREDICT on two distinct environments: a gridworld setting and a new text-domain environment (PLUME). PREDICT more accurately infers nuanced human preferences improving over existing baselines by 66.2\% (gridworld environment) and 41.0\% (PLUME). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_06273 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | PREDICT: Preference Reasoning by Evaluating Decomposed preferences Inferred from Candidate Trajectories Aroca-Ouellette, Stephane Mackraz, Natalie Theobald, Barry-John Metcalf, Katherine Artificial Intelligence Human-Computer Interaction Accommodating human preferences is essential for creating AI agents that deliver personalized and effective interactions. Recent work has shown the potential for LLMs to infer preferences from user interactions, but they often produce broad and generic preferences, failing to capture the unique and individualized nature of human preferences. This paper introduces PREDICT, a method designed to enhance the precision and adaptability of inferring preferences. PREDICT incorporates three key elements: (1) iterative refinement of inferred preferences, (2) decomposition of preferences into constituent components, and (3) validation of preferences across multiple trajectories. We evaluate PREDICT on two distinct environments: a gridworld setting and a new text-domain environment (PLUME). PREDICT more accurately infers nuanced human preferences improving over existing baselines by 66.2\% (gridworld environment) and 41.0\% (PLUME). |
| title | PREDICT: Preference Reasoning by Evaluating Decomposed preferences Inferred from Candidate Trajectories |
| topic | Artificial Intelligence Human-Computer Interaction |
| url | https://arxiv.org/abs/2410.06273 |