Probabilistic Satisfaction of Temporal Logic Constraints in Reinforcement Learning via Adaptive Policy-Switching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Xiaoshan, Yüksel, Sadık Bera, Yazıcıoğlu, Yasin, Aksaray, Derya
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929608200290304
author Lin, Xiaoshan
Yüksel, Sadık Bera
Yazıcıoğlu, Yasin
Aksaray, Derya
author_facet Lin, Xiaoshan
Yüksel, Sadık Bera
Yazıcıoğlu, Yasin
Aksaray, Derya
contents Constrained Reinforcement Learning (CRL) is a subset of machine learning that introduces constraints into the traditional reinforcement learning (RL) framework. Unlike conventional RL which aims solely to maximize cumulative rewards, CRL incorporates additional constraints that represent specific mission requirements or limitations that the agent must comply with during the learning process. In this paper, we address a type of CRL problem where an agent aims to learn the optimal policy to maximize reward while ensuring a desired level of temporal logic constraint satisfaction throughout the learning process. We propose a novel framework that relies on switching between pure learning (reward maximization) and constraint satisfaction. This framework estimates the probability of constraint satisfaction based on earlier trials and properly adjusts the probability of switching between learning and constraint satisfaction policies. We theoretically validate the correctness of the proposed algorithm and demonstrate its performance through comprehensive simulations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08022
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Probabilistic Satisfaction of Temporal Logic Constraints in Reinforcement Learning via Adaptive Policy-Switching
Lin, Xiaoshan
Yüksel, Sadık Bera
Yazıcıoğlu, Yasin
Aksaray, Derya
Artificial Intelligence
Robotics
Systems and Control
Constrained Reinforcement Learning (CRL) is a subset of machine learning that introduces constraints into the traditional reinforcement learning (RL) framework. Unlike conventional RL which aims solely to maximize cumulative rewards, CRL incorporates additional constraints that represent specific mission requirements or limitations that the agent must comply with during the learning process. In this paper, we address a type of CRL problem where an agent aims to learn the optimal policy to maximize reward while ensuring a desired level of temporal logic constraint satisfaction throughout the learning process. We propose a novel framework that relies on switching between pure learning (reward maximization) and constraint satisfaction. This framework estimates the probability of constraint satisfaction based on earlier trials and properly adjusts the probability of switching between learning and constraint satisfaction policies. We theoretically validate the correctness of the proposed algorithm and demonstrate its performance through comprehensive simulations.
title Probabilistic Satisfaction of Temporal Logic Constraints in Reinforcement Learning via Adaptive Policy-Switching
topic Artificial Intelligence
Robotics
Systems and Control
url https://arxiv.org/abs/2410.08022