Surprise-Adaptive Intrinsic Motivation for Unsupervised Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hugessen, Adriana, Castanyer, Roger Creus, Mohamed, Faisal, Berseth, Glen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910568163573760
author Hugessen, Adriana
Castanyer, Roger Creus
Mohamed, Faisal
Berseth, Glen
author_facet Hugessen, Adriana
Castanyer, Roger Creus
Mohamed, Faisal
Berseth, Glen
contents Both entropy-minimizing and entropy-maximizing (curiosity) objectives for unsupervised reinforcement learning (RL) have been shown to be effective in different environments, depending on the environment's level of natural entropy. However, neither method alone results in an agent that will consistently learn intelligent behavior across environments. In an effort to find a single entropy-based method that will encourage emergent behaviors in any environment, we propose an agent that can adapt its objective online, depending on the entropy conditions by framing the choice as a multi-armed bandit problem. We devise a novel intrinsic feedback signal for the bandit, which captures the agent's ability to control the entropy in its environment. We demonstrate that such agents can learn to control entropy and exhibit emergent behaviors in both high- and low-entropy regimes and can learn skillful behaviors in benchmark tasks. Videos of the trained agents and summarized findings can be found on our project page https://sites.google.com/view/surprise-adaptive-agents
format Preprint
id arxiv_https___arxiv_org_abs_2405_17243
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Surprise-Adaptive Intrinsic Motivation for Unsupervised Reinforcement Learning
Hugessen, Adriana
Castanyer, Roger Creus
Mohamed, Faisal
Berseth, Glen
Machine Learning
Artificial Intelligence
Both entropy-minimizing and entropy-maximizing (curiosity) objectives for unsupervised reinforcement learning (RL) have been shown to be effective in different environments, depending on the environment's level of natural entropy. However, neither method alone results in an agent that will consistently learn intelligent behavior across environments. In an effort to find a single entropy-based method that will encourage emergent behaviors in any environment, we propose an agent that can adapt its objective online, depending on the entropy conditions by framing the choice as a multi-armed bandit problem. We devise a novel intrinsic feedback signal for the bandit, which captures the agent's ability to control the entropy in its environment. We demonstrate that such agents can learn to control entropy and exhibit emergent behaviors in both high- and low-entropy regimes and can learn skillful behaviors in benchmark tasks. Videos of the trained agents and summarized findings can be found on our project page https://sites.google.com/view/surprise-adaptive-agents
title Surprise-Adaptive Intrinsic Motivation for Unsupervised Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.17243