Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Shao, Daqian, Buening, Thomas Kleine, Kwiatkowska, Marta
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914292290289664
author Shao, Daqian
Buening, Thomas Kleine
Kwiatkowska, Marta
author_facet Shao, Daqian
Buening, Thomas Kleine
Kwiatkowska, Marta
contents We propose a general framework for causal Imitation Learning (IL) with hidden confounders, which subsumes several existing settings. Our framework accounts for two types of hidden confounders: (a) variables observed by the expert but not by the imitator, and (b) confounding noise hidden from both. By leveraging trajectory histories as instruments, we reformulate causal IL in our framework into a Conditional Moment Restriction (CMR) problem. We propose DML-IL, an algorithm that solves this CMR problem via instrumental variable regression, and upper bound its imitation gap. Empirical evaluation on continuous state-action environments, including Mujoco tasks, demonstrates that DML-IL outperforms existing causal IL baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07656
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding
Shao, Daqian
Buening, Thomas Kleine
Kwiatkowska, Marta
Machine Learning
Artificial Intelligence
We propose a general framework for causal Imitation Learning (IL) with hidden confounders, which subsumes several existing settings. Our framework accounts for two types of hidden confounders: (a) variables observed by the expert but not by the imitator, and (b) confounding noise hidden from both. By leveraging trajectory histories as instruments, we reformulate causal IL in our framework into a Conditional Moment Restriction (CMR) problem. We propose DML-IL, an algorithm that solves this CMR problem via instrumental variable regression, and upper bound its imitation gap. Empirical evaluation on continuous state-action environments, including Mujoco tasks, demonstrates that DML-IL outperforms existing causal IL baselines.
title Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.07656