Enhancing RL Safety with Counterfactual LLM Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gross, Dennis, Spieker, Helge
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912030570577920
author Gross, Dennis
Spieker, Helge
author_facet Gross, Dennis
Spieker, Helge
contents Reinforcement learning (RL) policies may exhibit unsafe behavior and are hard to explain. We use counterfactual large language model reasoning to enhance RL policy safety post-training. We show that our approach improves and helps to explain the RL policy safety.
format Preprint
id arxiv_https___arxiv_org_abs_2409_10188
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing RL Safety with Counterfactual LLM Reasoning
Gross, Dennis
Spieker, Helge
Machine Learning
Reinforcement learning (RL) policies may exhibit unsafe behavior and are hard to explain. We use counterfactual large language model reasoning to enhance RL policy safety post-training. We show that our approach improves and helps to explain the RL policy safety.
title Enhancing RL Safety with Counterfactual LLM Reasoning
topic Machine Learning
url https://arxiv.org/abs/2409.10188