Demonstration Sidetracks: Categorizing Systematic Non-Optimality in Human Demonstrations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fang, Shijie, Yu, Hang, Fang, Qidi, Aronson, Reuben M., Short, Elaine S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917149756358656
author Fang, Shijie
Yu, Hang
Fang, Qidi
Aronson, Reuben M.
Short, Elaine S.
author_facet Fang, Shijie
Yu, Hang
Fang, Qidi
Aronson, Reuben M.
Short, Elaine S.
contents Learning from Demonstration (LfD) is a popular approach for robots to acquire new skills, but most LfD methods suffer from imperfections in human demonstrations. Prior work typically treats these suboptimalities as random noise. In this paper we study non-optimal behaviors in non-expert demonstrations and show that they are systematic, forming what we call demonstration sidetracks. Using a public space study with 40 participants performing a long-horizon robot task, we recreated the setup in simulation and annotated all demonstrations. We identify four types of sidetracks (Exploration, Mistake, Alignment, Pause) and one control pattern (one-dimension control). Sidetracks appear frequently across participants, and their temporal and spatial distribution is tied to task context. We also find that users' control patterns depend on the control interface. These insights point to the need for better models of suboptimal demonstrations to improve LfD algorithms and bridge the gap between lab training and real-world deployment. All demonstrations, infrastructure, and annotations are available at https://github.com/AABL-Lab/Human-Demonstration-Sidetracks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11262
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Demonstration Sidetracks: Categorizing Systematic Non-Optimality in Human Demonstrations
Fang, Shijie
Yu, Hang
Fang, Qidi
Aronson, Reuben M.
Short, Elaine S.
Robotics
Machine Learning
Learning from Demonstration (LfD) is a popular approach for robots to acquire new skills, but most LfD methods suffer from imperfections in human demonstrations. Prior work typically treats these suboptimalities as random noise. In this paper we study non-optimal behaviors in non-expert demonstrations and show that they are systematic, forming what we call demonstration sidetracks. Using a public space study with 40 participants performing a long-horizon robot task, we recreated the setup in simulation and annotated all demonstrations. We identify four types of sidetracks (Exploration, Mistake, Alignment, Pause) and one control pattern (one-dimension control). Sidetracks appear frequently across participants, and their temporal and spatial distribution is tied to task context. We also find that users' control patterns depend on the control interface. These insights point to the need for better models of suboptimal demonstrations to improve LfD algorithms and bridge the gap between lab training and real-world deployment. All demonstrations, infrastructure, and annotations are available at https://github.com/AABL-Lab/Human-Demonstration-Sidetracks.
title Demonstration Sidetracks: Categorizing Systematic Non-Optimality in Human Demonstrations
topic Robotics
Machine Learning
url https://arxiv.org/abs/2506.11262