TubeDAgger: Reducing the Number of Expert Interventions with Stochastic Reach-Tubes

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lemmel, Julian, Kranzl, Manuel, Lamine, Adam, Neubauer, Philipp, Grosu, Radu, Neubauer, Sophie A.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914070438871040
author Lemmel, Julian
Kranzl, Manuel
Lamine, Adam
Neubauer, Philipp
Grosu, Radu
Neubauer, Sophie A.
author_facet Lemmel, Julian
Kranzl, Manuel
Lamine, Adam
Neubauer, Philipp
Grosu, Radu
Neubauer, Sophie A.
contents Interactive Imitation Learning deals with training a novice policy from expert demonstrations in an online fashion. The established DAgger algorithm trains a robust novice policy by alternating between interacting with the environment and retraining of the network. Many variants thereof exist, that differ in the method of discerning whether to allow the novice to act or return control to the expert. We propose the use of stochastic reachtubes - common in verification of dynamical systems - as a novel method for estimating the necessity of expert intervention. Our approach does not require fine-tuning of decision thresholds per environment and effectively reduces the number of expert interventions, especially when compared with related approaches that make use of a doubt classification model.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00906
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TubeDAgger: Reducing the Number of Expert Interventions with Stochastic Reach-Tubes
Lemmel, Julian
Kranzl, Manuel
Lamine, Adam
Neubauer, Philipp
Grosu, Radu
Neubauer, Sophie A.
Systems and Control
Artificial Intelligence
Machine Learning
Interactive Imitation Learning deals with training a novice policy from expert demonstrations in an online fashion. The established DAgger algorithm trains a robust novice policy by alternating between interacting with the environment and retraining of the network. Many variants thereof exist, that differ in the method of discerning whether to allow the novice to act or return control to the expert. We propose the use of stochastic reachtubes - common in verification of dynamical systems - as a novel method for estimating the necessity of expert intervention. Our approach does not require fine-tuning of decision thresholds per environment and effectively reduces the number of expert interventions, especially when compared with related approaches that make use of a doubt classification model.
title TubeDAgger: Reducing the Number of Expert Interventions with Stochastic Reach-Tubes
topic Systems and Control
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.00906