Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866914292290289664 |
|---|---|
| author | Shao, Daqian Buening, Thomas Kleine Kwiatkowska, Marta |
| author_facet | Shao, Daqian Buening, Thomas Kleine Kwiatkowska, Marta |
| contents | We propose a general framework for causal Imitation Learning (IL) with hidden confounders, which subsumes several existing settings. Our framework accounts for two types of hidden confounders: (a) variables observed by the expert but not by the imitator, and (b) confounding noise hidden from both. By leveraging trajectory histories as instruments, we reformulate causal IL in our framework into a Conditional Moment Restriction (CMR) problem. We propose DML-IL, an algorithm that solves this CMR problem via instrumental variable regression, and upper bound its imitation gap. Empirical evaluation on continuous state-action environments, including Mujoco tasks, demonstrates that DML-IL outperforms existing causal IL baselines. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_07656 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding Shao, Daqian Buening, Thomas Kleine Kwiatkowska, Marta Machine Learning Artificial Intelligence We propose a general framework for causal Imitation Learning (IL) with hidden confounders, which subsumes several existing settings. Our framework accounts for two types of hidden confounders: (a) variables observed by the expert but not by the imitator, and (b) confounding noise hidden from both. By leveraging trajectory histories as instruments, we reformulate causal IL in our framework into a Conditional Moment Restriction (CMR) problem. We propose DML-IL, an algorithm that solves this CMR problem via instrumental variable regression, and upper bound its imitation gap. Empirical evaluation on continuous state-action environments, including Mujoco tasks, demonstrates that DML-IL outperforms existing causal IL baselines. |
| title | Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2502.07656 |