Action abstractions for amortized sampling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Boussif, Oussama, Ezzine, Léna Néhale, Viviano, Joseph D, Koziarski, Michał, Jain, Moksh, Malkin, Nikolay, Bengio, Emmanuel, Assouel, Rim, Bengio, Yoshua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912078586970112
author Boussif, Oussama
Ezzine, Léna Néhale
Viviano, Joseph D
Koziarski, Michał
Jain, Moksh
Malkin, Nikolay
Bengio, Emmanuel
Assouel, Rim
Bengio, Yoshua
author_facet Boussif, Oussama
Ezzine, Léna Néhale
Viviano, Joseph D
Koziarski, Michał
Jain, Moksh
Malkin, Nikolay
Bengio, Emmanuel
Assouel, Rim
Bengio, Yoshua
contents As trajectories sampled by policies used by reinforcement learning (RL) and generative flow networks (GFlowNets) grow longer, credit assignment and exploration become more challenging, and the long planning horizon hinders mode discovery and generalization. The challenge is particularly pronounced in entropy-seeking RL methods, such as generative flow networks, where the agent must learn to sample from a structured distribution and discover multiple high-reward states, each of which take many steps to reach. To tackle this challenge, we propose an approach to incorporate the discovery of action abstractions, or high-level actions, into the policy optimization process. Our approach involves iteratively extracting action subsequences commonly used across many high-reward trajectories and `chunking' them into a single action that is added to the action space. In empirical evaluation on synthetic and real-world environments, our approach demonstrates improved sample efficiency performance in discovering diverse high-reward objects, especially on harder exploration problems. We also observe that the abstracted high-order actions are interpretable, capturing the latent structure of the reward landscape of the action space. This work provides a cognitively motivated approach to action abstraction in RL and is the first demonstration of hierarchical planning in amortized sequential sampling.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15184
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Action abstractions for amortized sampling
Boussif, Oussama
Ezzine, Léna Néhale
Viviano, Joseph D
Koziarski, Michał
Jain, Moksh
Malkin, Nikolay
Bengio, Emmanuel
Assouel, Rim
Bengio, Yoshua
Machine Learning
Artificial Intelligence
As trajectories sampled by policies used by reinforcement learning (RL) and generative flow networks (GFlowNets) grow longer, credit assignment and exploration become more challenging, and the long planning horizon hinders mode discovery and generalization. The challenge is particularly pronounced in entropy-seeking RL methods, such as generative flow networks, where the agent must learn to sample from a structured distribution and discover multiple high-reward states, each of which take many steps to reach. To tackle this challenge, we propose an approach to incorporate the discovery of action abstractions, or high-level actions, into the policy optimization process. Our approach involves iteratively extracting action subsequences commonly used across many high-reward trajectories and `chunking' them into a single action that is added to the action space. In empirical evaluation on synthetic and real-world environments, our approach demonstrates improved sample efficiency performance in discovering diverse high-reward objects, especially on harder exploration problems. We also observe that the abstracted high-order actions are interpretable, capturing the latent structure of the reward landscape of the action space. This work provides a cognitively motivated approach to action abstraction in RL and is the first demonstration of hierarchical planning in amortized sequential sampling.
title Action abstractions for amortized sampling
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.15184