When Actions Disappear: Adversarial Action Removal in Self-Play Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autore principale: Kujur, Arahan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914571078336512
author Kujur, Arahan
author_facet Kujur, Arahan
contents We study adversarial action masking in self-play reinforcement learning: an attacker selectively removes legal actions from a victim's action set. Unlike observation or action perturbations, removal eliminates decision options before the agent acts. Across poker games scaling from 6 to 5,531 information states and two non-poker domains, learned masking causes substantially more damage than random masking and learned perturbation baselines. The attack persists across Q-learning, PPO, NFSP, neural NFSP, and DQN victims; transfers across agents; is amplified by self-play; and shows no recovery under extended masked training. Mechanistically, the adversary targets high-value decision points, captured by reach-weighted contingent action capacity (CAC$_w$) and a value-weighted refinement CAC$_v$. These results identify action availability as a distinct robustness surface in self-play RL.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16312
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Actions Disappear: Adversarial Action Removal in Self-Play Reinforcement Learning
Kujur, Arahan
Machine Learning
Artificial Intelligence
I.2.6; I.2.11
We study adversarial action masking in self-play reinforcement learning: an attacker selectively removes legal actions from a victim's action set. Unlike observation or action perturbations, removal eliminates decision options before the agent acts. Across poker games scaling from 6 to 5,531 information states and two non-poker domains, learned masking causes substantially more damage than random masking and learned perturbation baselines. The attack persists across Q-learning, PPO, NFSP, neural NFSP, and DQN victims; transfers across agents; is amplified by self-play; and shows no recovery under extended masked training. Mechanistically, the adversary targets high-value decision points, captured by reach-weighted contingent action capacity (CAC$_w$) and a value-weighted refinement CAC$_v$. These results identify action availability as a distinct robustness surface in self-play RL.
title When Actions Disappear: Adversarial Action Removal in Self-Play Reinforcement Learning
topic Machine Learning
Artificial Intelligence
I.2.6; I.2.11
url https://arxiv.org/abs/2605.16312