Sparse Masked Attention Policies for Reliable Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Horsch, Caroline, Engwegen, Laurens, Weltevrede, Max, Spaan, Matthijs T. J., Böhmer, Wendelin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2025)
by: Weltevrede, Max, et al.
Published: (2025)
Explore-Go: Leveraging Exploration for Generalisation in Deep Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
Generalisation to unseen topologies: Towards control of biological neural network activity
by: Engwegen, Laurens, et al.
Published: (2024)
by: Engwegen, Laurens, et al.
Published: (2024)
Universal Value-Function Uncertainties
by: Zanger, Moritz A., et al.
Published: (2025)
by: Zanger, Moritz A., et al.
Published: (2025)
Diverse Projection Ensembles for Distributional Reinforcement Learning
by: Zanger, Moritz A., et al.
Published: (2023)
by: Zanger, Moritz A., et al.
Published: (2023)
Epistemic Monte Carlo Tree Search
by: Oren, Yaniv, et al.
Published: (2022)
by: Oren, Yaniv, et al.
Published: (2022)
Modular Recurrence in Contextual MDPs for Universal Morphology Control
by: Engwegen, Laurens, et al.
Published: (2025)
by: Engwegen, Laurens, et al.
Published: (2025)
Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model
by: Zanger, Moritz A., et al.
Published: (2025)
by: Zanger, Moritz A., et al.
Published: (2025)
PMCTS: Particle Monte Carlo Tree Search for Principled Parallelized Inference Time Scaling
by: Oren, Yaniv, et al.
Published: (2026)
by: Oren, Yaniv, et al.
Published: (2026)
On the Equivalence of Random Network Distillation, Deep Ensembles, and Bayesian Inference
by: Zanger, Moritz A., et al.
Published: (2026)
by: Zanger, Moritz A., et al.
Published: (2026)
Twice Sequential Monte Carlo for Tree Search
by: Oren, Yaniv, et al.
Published: (2025)
by: Oren, Yaniv, et al.
Published: (2025)
Bad Habits: Policy Confounding and Out-of-Trajectory Generalization in RL
by: Suau, Miguel, et al.
Published: (2023)
by: Suau, Miguel, et al.
Published: (2023)
Value Improved Actor Critic Algorithms
by: Oren, Yaniv, et al.
Published: (2024)
by: Oren, Yaniv, et al.
Published: (2024)
Improving Robustness of AlphaZero Algorithms to Test-Time Environment Changes
by: Tamassia, Isidoro, et al.
Published: (2025)
by: Tamassia, Isidoro, et al.
Published: (2025)
TransZero: Parallel Tree Expansion in MuZero using Transformer Networks
by: Malmsten, Emil, et al.
Published: (2025)
by: Malmsten, Emil, et al.
Published: (2025)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
by: Li, Guopeng, et al.
Published: (2026)
by: Li, Guopeng, et al.
Published: (2026)
When Do Off-Policy and On-Policy Policy Gradient Methods Align?
by: Mambelli, Davide, et al.
Published: (2024)
by: Mambelli, Davide, et al.
Published: (2024)
Trust-Region Twisted Policy Improvement
by: de Vries, Joery A., et al.
Published: (2025)
by: de Vries, Joery A., et al.
Published: (2025)
To the Max: Reinventing Reward in Reinforcement Learning
by: Veviurko, Grigorii, et al.
Published: (2024)
by: Veviurko, Grigorii, et al.
Published: (2024)
AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play
by: Murgoci, Vlad, et al.
Published: (2026)
by: Murgoci, Vlad, et al.
Published: (2026)
You Shall Pass: Dealing with the Zero-Gradient Problem in Predict and Optimize for Convex Optimization
by: Veviurko, Grigorii, et al.
Published: (2023)
by: Veviurko, Grigorii, et al.
Published: (2023)
Positive Experience Reflection for Agents in Interactive Text Environments
by: Lippmann, Philip, et al.
Published: (2024)
by: Lippmann, Philip, et al.
Published: (2024)
EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control
by: Evers, Thomas, et al.
Published: (2026)
by: Evers, Thomas, et al.
Published: (2026)
SEA: Sparse Linear Attention with Estimated Attention Mask
by: Lee, Heejun, et al.
Published: (2023)
by: Lee, Heejun, et al.
Published: (2023)
Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning
by: van der Vaart, Pascal R., et al.
Published: (2025)
by: van der Vaart, Pascal R., et al.
Published: (2025)
Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks
by: de Vries, Joery A., et al.
Published: (2025)
by: de Vries, Joery A., et al.
Published: (2025)
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)
by: Shi, Jingze, et al.
Published: (2025)
Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems
by: Suau, Miguel, et al.
Published: (2022)
by: Suau, Miguel, et al.
Published: (2022)
Reinforcement Learning by Guided Safe Exploration
by: Yang, Qisong, et al.
Published: (2023)
by: Yang, Qisong, et al.
Published: (2023)
TamedPUMA: safe and stable imitation learning with geometric fabrics
by: Bakker, Saray, et al.
Published: (2025)
by: Bakker, Saray, et al.
Published: (2025)
RecBayes: Recurrent Bayesian Ad Hoc Teamwork in Large Partially Observable Domains
by: Ribeiro, João G., et al.
Published: (2025)
by: Ribeiro, João G., et al.
Published: (2025)
SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale
by: Zimmer, Max, et al.
Published: (2025)
by: Zimmer, Max, et al.
Published: (2025)
Deep Gaussian Process Proximal Policy Optimization
by: van der Lende, Matthijs, et al.
Published: (2025)
by: van der Lende, Matthijs, et al.
Published: (2025)
FlashMask: Efficient and Rich Mask Extension of FlashAttention
by: Wang, Guoxia, et al.
Published: (2024)
by: Wang, Guoxia, et al.
Published: (2024)
VariBASed: Variational Bayes-Adaptive Sequential Monte-Carlo Planning for Deep Reinforcement Learning
by: de Vries, Joery A., et al.
Published: (2026)
by: de Vries, Joery A., et al.
Published: (2026)
Pessimistic Iterative Planning with RNNs for Robust POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2024)
by: Galesloot, Maris F. L., et al.
Published: (2024)
Machine learning and high dimensional vector search
by: Douze, Matthijs
Published: (2025)
by: Douze, Matthijs
Published: (2025)
Are Sparse Autoencoder Benchmarks Reliable?
by: Chanin, David
Published: (2026)
by: Chanin, David
Published: (2026)
On the Role of Attention Masks and LayerNorm in Transformers
by: Wu, Xinyi, et al.
Published: (2024)
by: Wu, Xinyi, et al.
Published: (2024)
Similar Items
-
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
by: Weltevrede, Max, et al.
Published: (2024) -
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2025) -
Explore-Go: Leveraging Exploration for Generalisation in Deep Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2024) -
Generalisation to unseen topologies: Towards control of biological neural network activity
by: Engwegen, Laurens, et al.
Published: (2024) -
Universal Value-Function Uncertainties
by: Zanger, Moritz A., et al.
Published: (2025)