Sparse Masked Attention Policies for Reliable Generalization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Horsch, Caroline, Engwegen, Laurens, Weltevrede, Max, Spaan, Matthijs T. J., Böhmer, Wendelin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911463073906688
author Horsch, Caroline
Engwegen, Laurens
Weltevrede, Max
Spaan, Matthijs T. J.
Böhmer, Wendelin
author_facet Horsch, Caroline
Engwegen, Laurens
Weltevrede, Max
Spaan, Matthijs T. J.
Böhmer, Wendelin
contents In reinforcement learning, abstraction methods that remove unnecessary information from the observation are commonly used to learn policies which generalize better to unseen tasks. However, these methods often overlook a crucial weakness: the function which extracts the reduced-information representation has unknown generalization ability in unseen observations. In this paper, we address this problem by presenting an information removal method which more reliably generalizes to new states. We accomplish this by using a learned masking function which operates on, and is integrated with, the attention weights within an attention-based policy network. We demonstrate that our method significantly improves policy generalization to unseen tasks in the Procgen benchmark compared to standard PPO and masking approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19956
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Sparse Masked Attention Policies for Reliable Generalization
Horsch, Caroline
Engwegen, Laurens
Weltevrede, Max
Spaan, Matthijs T. J.
Böhmer, Wendelin
Machine Learning
In reinforcement learning, abstraction methods that remove unnecessary information from the observation are commonly used to learn policies which generalize better to unseen tasks. However, these methods often overlook a crucial weakness: the function which extracts the reduced-information representation has unknown generalization ability in unseen observations. In this paper, we address this problem by presenting an information removal method which more reliably generalizes to new states. We accomplish this by using a learned masking function which operates on, and is integrated with, the attention weights within an attention-based policy network. We demonstrate that our method significantly improves policy generalization to unseen tasks in the Procgen benchmark compared to standard PPO and masking approaches.
title Sparse Masked Attention Policies for Reliable Generalization
topic Machine Learning
url https://arxiv.org/abs/2602.19956