A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Kujur, Arahan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914571084627968
author Kujur, Arahan
author_facet Kujur, Arahan
contents We show that a threshold in decision capacity determines whether self-play reinforcement learning agents collapse under asymmetric rule perturbations. Across poker variants, matrix games, a dice game, and multiple learning algorithms, eliminating all positive-reach contingent decisions causes rapid convergence to a deterministic exploitation attractor, a fixed point at near-maximal loss. Preserving even a single positive-reach contingent decision point prevents this collapse. A frozen baseline and fixed-opponent control confirm that the mechanism is co-adaptation under constraint, not the perturbation itself. The phenomenon is timing-invariant, fully reversible upon action restoration, and intensifies under function approximation. These results establish a sharp threshold at zero reach-weighted contingent action capacity, with severity scaling continuously via reach-weighted capacity in the tested domains.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16315
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning
Kujur, Arahan
Machine Learning
Artificial Intelligence
91A05, 91A26, 68T05, 68T42
I.2.6; I.2.8; F.2.2
We show that a threshold in decision capacity determines whether self-play reinforcement learning agents collapse under asymmetric rule perturbations. Across poker variants, matrix games, a dice game, and multiple learning algorithms, eliminating all positive-reach contingent decisions causes rapid convergence to a deterministic exploitation attractor, a fixed point at near-maximal loss. Preserving even a single positive-reach contingent decision point prevents this collapse. A frozen baseline and fixed-opponent control confirm that the mechanism is co-adaptation under constraint, not the perturbation itself. The phenomenon is timing-invariant, fully reversible upon action restoration, and intensifies under function approximation. These results establish a sharp threshold at zero reach-weighted contingent action capacity, with severity scaling continuously via reach-weighted capacity in the tested domains.
title A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning
topic Machine Learning
Artificial Intelligence
91A05, 91A26, 68T05, 68T42
I.2.6; I.2.8; F.2.2
url https://arxiv.org/abs/2605.16315