Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Role of Bellman Constraints

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Tian, Wang, Chenyang, Zhai, Xiaochen, Li, Ziniu, Li, Yi-Chen, Yu, Yang
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911540928577536
author Xu, Tian
Wang, Chenyang
Zhai, Xiaochen
Li, Ziniu
Li, Yi-Chen
Yu, Yang
author_facet Xu, Tian
Wang, Chenyang
Zhai, Xiaochen
Li, Ziniu
Li, Yi-Chen
Yu, Yang
contents Adversarial imitation learning (AIL) achieves high-quality imitation by mitigating compounding errors in behavioral cloning (BC), but often exhibits training instability due to adversarial optimization. To avoid this issue, a class of non-adversarial Q-based imitation learning (IL) methods, represented by IQ-Learn, has emerged and is widely believed to outperform BC by leveraging online environment interactions. However, this paper revisits IQ-Learn and demonstrates that it provably reduces to BC and suffers from an imitation gap lower bound with quadratic dependence on horizon, therefore still suffering from compounding errors. Theoretical analysis reveals that, despite using online interactions, IQ-Learn uniformly suppresses the Q-values for all actions on states uncovered by demonstrations, thereby failing to generalize. To address this limitation, we introduce a primal-dual framework for distribution matching, yielding a new Q-based IL method, Dual Q-DM. The key mechanism in Dual Q-DM is incorporating Bellman constraints to propagate high Q-values from visited states to unvisited ones, thereby achieving generalization beyond demonstrations. We prove that Dual Q-DM is equivalent to AIL and can recover expert actions beyond demonstrations, thereby mitigating compounding errors. To the best of our knowledge, Dual Q-DM is the first non-adversarial IL method that is theoretically guaranteed to eliminate compounding errors. Experimental results further corroborate our theoretical results.
format Preprint
id arxiv_https___arxiv_org_abs_2603_22713
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Role of Bellman Constraints
Xu, Tian
Wang, Chenyang
Zhai, Xiaochen
Li, Ziniu
Li, Yi-Chen
Yu, Yang
Machine Learning
Adversarial imitation learning (AIL) achieves high-quality imitation by mitigating compounding errors in behavioral cloning (BC), but often exhibits training instability due to adversarial optimization. To avoid this issue, a class of non-adversarial Q-based imitation learning (IL) methods, represented by IQ-Learn, has emerged and is widely believed to outperform BC by leveraging online environment interactions. However, this paper revisits IQ-Learn and demonstrates that it provably reduces to BC and suffers from an imitation gap lower bound with quadratic dependence on horizon, therefore still suffering from compounding errors. Theoretical analysis reveals that, despite using online interactions, IQ-Learn uniformly suppresses the Q-values for all actions on states uncovered by demonstrations, thereby failing to generalize. To address this limitation, we introduce a primal-dual framework for distribution matching, yielding a new Q-based IL method, Dual Q-DM. The key mechanism in Dual Q-DM is incorporating Bellman constraints to propagate high Q-values from visited states to unvisited ones, thereby achieving generalization beyond demonstrations. We prove that Dual Q-DM is equivalent to AIL and can recover expert actions beyond demonstrations, thereby mitigating compounding errors. To the best of our knowledge, Dual Q-DM is the first non-adversarial IL method that is theoretically guaranteed to eliminate compounding errors. Experimental results further corroborate our theoretical results.
title Non-Adversarial Imitation Learning Provably Free of Compounding Errors: The Role of Bellman Constraints
topic Machine Learning
url https://arxiv.org/abs/2603.22713