Learning safe, constrained policies via imitation learning: Connection to Probabilistic Inference and a Naive Algorithm

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Papadopoulos, George, Vouros, George A.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913933845069824
author Papadopoulos, George
Vouros, George A.
author_facet Papadopoulos, George
Vouros, George A.
contents This article introduces an imitation learning method for learning maximum entropy policies that comply with constraints demonstrated by expert trajectories executing a task. The formulation of the method takes advantage of results connecting performance to bounds for the KL-divergence between demonstrated and learned policies, and its objective is rigorously justified through a connection to a probabilistic inference framework for reinforcement learning, incorporating the reinforcement learning objective and the objective to abide by constraints in an entropy maximization setting. The proposed algorithm optimizes the learning objective with dual gradient descent, supporting effective and stable training. Experiments show that the proposed method can learn effective policy models for constraints-abiding behaviour, in settings with multiple constraints of different types, accommodating different modalities of demonstrated behaviour, and with abilities to generalize.
format Preprint
id arxiv_https___arxiv_org_abs_2507_06780
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning safe, constrained policies via imitation learning: Connection to Probabilistic Inference and a Naive Algorithm
Papadopoulos, George
Vouros, George A.
Machine Learning
Multiagent Systems
This article introduces an imitation learning method for learning maximum entropy policies that comply with constraints demonstrated by expert trajectories executing a task. The formulation of the method takes advantage of results connecting performance to bounds for the KL-divergence between demonstrated and learned policies, and its objective is rigorously justified through a connection to a probabilistic inference framework for reinforcement learning, incorporating the reinforcement learning objective and the objective to abide by constraints in an entropy maximization setting. The proposed algorithm optimizes the learning objective with dual gradient descent, supporting effective and stable training. Experiments show that the proposed method can learn effective policy models for constraints-abiding behaviour, in settings with multiple constraints of different types, accommodating different modalities of demonstrated behaviour, and with abilities to generalize.
title Learning safe, constrained policies via imitation learning: Connection to Probabilistic Inference and a Naive Algorithm
topic Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2507.06780