Contrastive Learning from Exploratory Actions: Leveraging Natural Interactions for Preference Elicitation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dennler, Nathaniel, Nikolaidis, Stefanos, Matarić, Maja
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912175519432704
author Dennler, Nathaniel
Nikolaidis, Stefanos
Matarić, Maja
author_facet Dennler, Nathaniel
Nikolaidis, Stefanos
Matarić, Maja
contents People have a variety of preferences for how robots behave. To understand and reason about these preferences, robots aim to learn a reward function that describes how aligned robot behaviors are with a user's preferences. Good representations of a robot's behavior can significantly reduce the time and effort required for a user to teach the robot their preferences. Specifying these representations -- what "features" of the robot's behavior matter to users -- remains a difficult problem; Features learned from raw data lack semantic meaning and features learned from user data require users to engage in tedious labeling processes. Our key insight is that users tasked with customizing a robot are intrinsically motivated to produce labels through exploratory search; they explore behaviors that they find interesting and ignore behaviors that are irrelevant. To harness this novel data source of exploratory actions, we propose contrastive learning from exploratory actions (CLEA) to learn trajectory features that are aligned with features that users care about. We learned CLEA features from exploratory actions users performed in an open-ended signal design activity (N=25) with a Kuri robot, and evaluated CLEA features through a second user study with a different set of users (N=42). CLEA features outperformed self-supervised features when eliciting user preferences over four metrics: completeness, simplicity, minimality, and explainability.
format Preprint
id arxiv_https___arxiv_org_abs_2501_01367
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Contrastive Learning from Exploratory Actions: Leveraging Natural Interactions for Preference Elicitation
Dennler, Nathaniel
Nikolaidis, Stefanos
Matarić, Maja
Robotics
Artificial Intelligence
Human-Computer Interaction
Machine Learning
People have a variety of preferences for how robots behave. To understand and reason about these preferences, robots aim to learn a reward function that describes how aligned robot behaviors are with a user's preferences. Good representations of a robot's behavior can significantly reduce the time and effort required for a user to teach the robot their preferences. Specifying these representations -- what "features" of the robot's behavior matter to users -- remains a difficult problem; Features learned from raw data lack semantic meaning and features learned from user data require users to engage in tedious labeling processes. Our key insight is that users tasked with customizing a robot are intrinsically motivated to produce labels through exploratory search; they explore behaviors that they find interesting and ignore behaviors that are irrelevant. To harness this novel data source of exploratory actions, we propose contrastive learning from exploratory actions (CLEA) to learn trajectory features that are aligned with features that users care about. We learned CLEA features from exploratory actions users performed in an open-ended signal design activity (N=25) with a Kuri robot, and evaluated CLEA features through a second user study with a different set of users (N=42). CLEA features outperformed self-supervised features when eliciting user preferences over four metrics: completeness, simplicity, minimality, and explainability.
title Contrastive Learning from Exploratory Actions: Leveraging Natural Interactions for Preference Elicitation
topic Robotics
Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2501.01367