Adaptive Inverse Reinforcement Learning with Online Off-Policy Data Collection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Yibei, Cao, Yuexin, Liu, Zhixin, Xie, Lihua
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917092154933248
author Li, Yibei
Cao, Yuexin
Liu, Zhixin
Xie, Lihua
author_facet Li, Yibei
Cao, Yuexin
Liu, Zhixin
Xie, Lihua
contents In this paper, the inverse reinforcement learning (IRL) problem is addressed to reconstruct the unknown cost function underlying an observed optimal policy in a model-free manner, whose online adaptation with completely off-policy system data still remains unclear in the literature. Without prior knowledge of the system model parameters, an adaptive and direct learning rule for the cost parameter is proposed using online off-policy system data, which only needs to satisfy the mild persistently exciting condition in the general data-driven paradigm. The adaptive and online IRL algorithm is achieved by designing full Nesterov-Todd (NT)-step primal-dual interior-point iterations. Despite solving a nonlinear and time-varying semi-definite program (SDP), the influence of system noise is rigorously analyzed, and the proposed online algorithm is shown to achieve a sublinear convergence. The proposed method is further generalized to nonlinear IRL based on differential dynamic programming. The gradient of the loss function is directly obtained via a backward pass, which eliminates the need to repeatedly solve forward RL problems as in conventional bi-level IRL frameworks. Finally, the efficiency and effectiveness of the proposed algorithms are demonstrated by numerical examples.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15171
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adaptive Inverse Reinforcement Learning with Online Off-Policy Data Collection
Li, Yibei
Cao, Yuexin
Liu, Zhixin
Xie, Lihua
Optimization and Control
In this paper, the inverse reinforcement learning (IRL) problem is addressed to reconstruct the unknown cost function underlying an observed optimal policy in a model-free manner, whose online adaptation with completely off-policy system data still remains unclear in the literature. Without prior knowledge of the system model parameters, an adaptive and direct learning rule for the cost parameter is proposed using online off-policy system data, which only needs to satisfy the mild persistently exciting condition in the general data-driven paradigm. The adaptive and online IRL algorithm is achieved by designing full Nesterov-Todd (NT)-step primal-dual interior-point iterations. Despite solving a nonlinear and time-varying semi-definite program (SDP), the influence of system noise is rigorously analyzed, and the proposed online algorithm is shown to achieve a sublinear convergence. The proposed method is further generalized to nonlinear IRL based on differential dynamic programming. The gradient of the loss function is directly obtained via a backward pass, which eliminates the need to repeatedly solve forward RL problems as in conventional bi-level IRL frameworks. Finally, the efficiency and effectiveness of the proposed algorithms are demonstrated by numerical examples.
title Adaptive Inverse Reinforcement Learning with Online Off-Policy Data Collection
topic Optimization and Control
url https://arxiv.org/abs/2511.15171