OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Burnwal, Returaj, Bhatt, Nirav Pravinbhai, Ravindran, Balaraman
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912897842544640
author Burnwal, Returaj
Bhatt, Nirav Pravinbhai
Ravindran, Balaraman
author_facet Burnwal, Returaj
Bhatt, Nirav Pravinbhai
Ravindran, Balaraman
contents This work addresses the problem of offline safe imitation learning (IL), where the goal is to learn safe and reward-maximizing policies from demonstrations that do not have per-timestep safety cost or reward information. In many real-world domains, online learning in the environment can be risky, and specifying accurate safety costs can be difficult. However, it is often feasible to collect trajectories that reflect undesirable or unsafe behavior, implicitly conveying what the agent should avoid. We refer to these as non-preferred trajectories. We propose a novel offline safe IL algorithm, OSIL, that infers safety from non-preferred demonstrations. We formulate safe policy learning as a Constrained Markov Decision Process (CMDP). Instead of relying on explicit safety cost and reward annotations, OSIL reformulates the CMDP problem by deriving a lower bound on reward maximizing objective and learning a cost model that estimates the likelihood of non-preferred behavior. Our approach allows agents to learn safe and reward-maximizing behavior entirely from offline demonstrations. We empirically demonstrate that our approach can learn safer policies that satisfy cost constraints without degrading the reward performance, thus outperforming several baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2602_11018
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories
Burnwal, Returaj
Bhatt, Nirav Pravinbhai
Ravindran, Balaraman
Machine Learning
Artificial Intelligence
This work addresses the problem of offline safe imitation learning (IL), where the goal is to learn safe and reward-maximizing policies from demonstrations that do not have per-timestep safety cost or reward information. In many real-world domains, online learning in the environment can be risky, and specifying accurate safety costs can be difficult. However, it is often feasible to collect trajectories that reflect undesirable or unsafe behavior, implicitly conveying what the agent should avoid. We refer to these as non-preferred trajectories. We propose a novel offline safe IL algorithm, OSIL, that infers safety from non-preferred demonstrations. We formulate safe policy learning as a Constrained Markov Decision Process (CMDP). Instead of relying on explicit safety cost and reward annotations, OSIL reformulates the CMDP problem by deriving a lower bound on reward maximizing objective and learning a cost model that estimates the likelihood of non-preferred behavior. Our approach allows agents to learn safe and reward-maximizing behavior entirely from offline demonstrations. We empirically demonstrate that our approach can learn safer policies that satisfy cost constraints without degrading the reward performance, thus outperforming several baselines.
title OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.11018