A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Author: | |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914571084627968 |
|---|---|
| author | Kujur, Arahan |
| author_facet | Kujur, Arahan |
| contents | We show that a threshold in decision capacity determines whether self-play reinforcement learning agents collapse under asymmetric rule perturbations. Across poker variants, matrix games, a dice game, and multiple learning algorithms, eliminating all positive-reach contingent decisions causes rapid convergence to a deterministic exploitation attractor, a fixed point at near-maximal loss. Preserving even a single positive-reach contingent decision point prevents this collapse. A frozen baseline and fixed-opponent control confirm that the mechanism is co-adaptation under constraint, not the perturbation itself. The phenomenon is timing-invariant, fully reversible upon action restoration, and intensifies under function approximation. These results establish a sharp threshold at zero reach-weighted contingent action capacity, with severity scaling continuously via reach-weighted capacity in the tested domains. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_16315 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning Kujur, Arahan Machine Learning Artificial Intelligence 91A05, 91A26, 68T05, 68T42 I.2.6; I.2.8; F.2.2 We show that a threshold in decision capacity determines whether self-play reinforcement learning agents collapse under asymmetric rule perturbations. Across poker variants, matrix games, a dice game, and multiple learning algorithms, eliminating all positive-reach contingent decisions causes rapid convergence to a deterministic exploitation attractor, a fixed point at near-maximal loss. Preserving even a single positive-reach contingent decision point prevents this collapse. A frozen baseline and fixed-opponent control confirm that the mechanism is co-adaptation under constraint, not the perturbation itself. The phenomenon is timing-invariant, fully reversible upon action restoration, and intensifies under function approximation. These results establish a sharp threshold at zero reach-weighted contingent action capacity, with severity scaling continuously via reach-weighted capacity in the tested domains. |
| title | A Structural Threshold in Decision Capacity Governs Collapse in Self-Play Reinforcement Learning |
| topic | Machine Learning Artificial Intelligence 91A05, 91A26, 68T05, 68T42 I.2.6; I.2.8; F.2.2 |
| url | https://arxiv.org/abs/2605.16315 |