Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuhar, Sachit, Cheng, Shuo, Chopra, Shivang, Bronars, Matthew, Xu, Danfei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912361662644224
author Kuhar, Sachit
Cheng, Shuo
Chopra, Shivang
Bronars, Matthew
Xu, Danfei
author_facet Kuhar, Sachit
Cheng, Shuo
Chopra, Shivang
Bronars, Matthew
Xu, Danfei
contents Practical Imitation Learning (IL) systems rely on large human demonstration datasets for successful policy learning. However, challenges lie in maintaining the quality of collected data and addressing the suboptimal nature of some demonstrations, which can compromise the overall dataset quality and hence the learning outcome. Furthermore, the intrinsic heterogeneity in human behavior can produce equally successful but disparate demonstrations, further exacerbating the challenge of discerning demonstration quality. To address these challenges, this paper introduces Learning to Discern (L2D), an offline imitation learning framework for learning from demonstrations with diverse quality and style. Given a small batch of demonstrations with sparse quality labels, we learn a latent representation for temporally embedded trajectory segments. Preference learning in this latent space trains a quality evaluator that generalizes to new demonstrators exhibiting different styles. Empirically, we show that L2D can effectively assess and learn from varying demonstrations, thereby leading to improved policy performance across a range of tasks in both simulations and on a physical robot.
format Preprint
id arxiv_https___arxiv_org_abs_2310_14196
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning
Kuhar, Sachit
Cheng, Shuo
Chopra, Shivang
Bronars, Matthew
Xu, Danfei
Robotics
Artificial Intelligence
Practical Imitation Learning (IL) systems rely on large human demonstration datasets for successful policy learning. However, challenges lie in maintaining the quality of collected data and addressing the suboptimal nature of some demonstrations, which can compromise the overall dataset quality and hence the learning outcome. Furthermore, the intrinsic heterogeneity in human behavior can produce equally successful but disparate demonstrations, further exacerbating the challenge of discerning demonstration quality. To address these challenges, this paper introduces Learning to Discern (L2D), an offline imitation learning framework for learning from demonstrations with diverse quality and style. Given a small batch of demonstrations with sparse quality labels, we learn a latent representation for temporally embedded trajectory segments. Preference learning in this latent space trains a quality evaluator that generalizes to new demonstrators exhibiting different styles. Empirically, we show that L2D can effectively assess and learn from varying demonstrations, thereby leading to improved policy performance across a range of tasks in both simulations and on a physical robot.
title Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2310.14196