Saved in:
Bibliographic Details
Main Authors: Low, Siow Meng, Kumar, Akshat
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.12557
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915253652029440
author Low, Siow Meng
Kumar, Akshat
author_facet Low, Siow Meng
Kumar, Akshat
contents In safe reinforcement learning (RL), auxiliary safety costs are used to align the agent to safe decision making. In practice, safety constraints, including cost functions and budgets, are unknown or hard to specify, as it requires anticipation of all possible unsafe behaviors. We therefore address a general setting where the true safety definition is unknown, and has to be learned from sparsely labeled data. Our key contributions are: first, we design a safety model that performs credit assignment to estimate each decision step's impact on the overall safety using a dataset of diverse trajectories and their corresponding binary safety labels (i.e., whether the corresponding trajectory is safe/unsafe). Second, we illustrate the architecture of our safety model to demonstrate its ability to learn a separate safety score for each timestep. Third, we reformulate the safe RL problem using the proposed safety model and derive an effective algorithm to optimize a safe yet rewarding policy. Finally, our empirical results corroborate our findings and show that this approach is effective in satisfying unknown safety definition, and scalable to various continuous control tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12557
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TraCeS: Trajectory Based Credit Assignment From Sparse Safety Feedback
Low, Siow Meng
Kumar, Akshat
Machine Learning
Artificial Intelligence
In safe reinforcement learning (RL), auxiliary safety costs are used to align the agent to safe decision making. In practice, safety constraints, including cost functions and budgets, are unknown or hard to specify, as it requires anticipation of all possible unsafe behaviors. We therefore address a general setting where the true safety definition is unknown, and has to be learned from sparsely labeled data. Our key contributions are: first, we design a safety model that performs credit assignment to estimate each decision step's impact on the overall safety using a dataset of diverse trajectories and their corresponding binary safety labels (i.e., whether the corresponding trajectory is safe/unsafe). Second, we illustrate the architecture of our safety model to demonstrate its ability to learn a separate safety score for each timestep. Third, we reformulate the safe RL problem using the proposed safety model and derive an effective algorithm to optimize a safe yet rewarding policy. Finally, our empirical results corroborate our findings and show that this approach is effective in satisfying unknown safety definition, and scalable to various continuous control tasks.
title TraCeS: Trajectory Based Credit Assignment From Sparse Safety Feedback
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.12557