Learning to Be Cautious

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mohammedalamen, Montaser, Morrill, Dustin, Sieusahai, Alexander, Satsangi, Yash, Bowling, Michael
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914087040974848
author Mohammedalamen, Montaser
Morrill, Dustin
Sieusahai, Alexander
Satsangi, Yash
Bowling, Michael
author_facet Mohammedalamen, Montaser
Morrill, Dustin
Sieusahai, Alexander
Satsangi, Yash
Bowling, Michael
contents A key challenge in the field of reinforcement learning is to develop agents that behave cautiously in novel situations. It is generally impossible to anticipate all situations that an autonomous system may face or what behavior would best avoid bad outcomes. An agent that can learn to be cautious would overcome this challenge by discovering for itself when and how to behave cautiously. In contrast, current approaches typically embed task-specific safety information or explicit cautious behaviors into the system, which is error-prone and imposes extra burdens on practitioners. In this paper, we present both a sequence of tasks where cautious behavior becomes increasingly non-obvious, as well as an algorithm to demonstrate that it is possible for a system to learn to be cautious. The essential features of our algorithm are that it characterizes reward function uncertainty without task-specific safety information and uses this uncertainty to construct a robust policy. Specifically, we construct robust policies with a k-of-N counterfactual regret minimization (CFR) subroutine given learned reward function uncertainty represented by a neural network ensemble. These policies exhibit caution in each of our tasks without any task-specific safety tuning. Our code is available at https://github.com/montaserFath/Learning-to-be-Cautious
format Preprint
id arxiv_https___arxiv_org_abs_2110_15907
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Learning to Be Cautious
Mohammedalamen, Montaser
Morrill, Dustin
Sieusahai, Alexander
Satsangi, Yash
Bowling, Michael
Artificial Intelligence
Machine Learning
A key challenge in the field of reinforcement learning is to develop agents that behave cautiously in novel situations. It is generally impossible to anticipate all situations that an autonomous system may face or what behavior would best avoid bad outcomes. An agent that can learn to be cautious would overcome this challenge by discovering for itself when and how to behave cautiously. In contrast, current approaches typically embed task-specific safety information or explicit cautious behaviors into the system, which is error-prone and imposes extra burdens on practitioners. In this paper, we present both a sequence of tasks where cautious behavior becomes increasingly non-obvious, as well as an algorithm to demonstrate that it is possible for a system to learn to be cautious. The essential features of our algorithm are that it characterizes reward function uncertainty without task-specific safety information and uses this uncertainty to construct a robust policy. Specifically, we construct robust policies with a k-of-N counterfactual regret minimization (CFR) subroutine given learned reward function uncertainty represented by a neural network ensemble. These policies exhibit caution in each of our tasks without any task-specific safety tuning. Our code is available at https://github.com/montaserFath/Learning-to-be-Cautious
title Learning to Be Cautious
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2110.15907