Commitment Floors for Tipping-Point Commons: Escaping Nash Traps in Multi-Agent Reinforcement Learning
Fuente:
Zenodo
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Recurso digital |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866901668273061888 |
|---|---|
| author | Author, Anonymous |
| author_facet | Author, Anonymous |
| contents | [Note 2026-05-18: This record's attached file name contains legacy internal naming that includes a venue label; the file name does not represent any active submission status, and the attached content is an anonymous preprint. File renaming is not permitted on published Zenodo records; a Zenodo Support request for either file rename or record withdrawal has been submitted.]<br><br><p><strong>Status:</strong> [redacted venue] [redacted venue] under double-blind review. Author identity anonymized.</p><p>We characterize cooperation failure in <strong>Tipping-Point Social Dilemmas</strong> (TPSDs) -- multi-agent settings where thresholded resource dynamics can trigger irreversible collapse. We identify the <strong>Nash Trap</strong>: a locally stable under-commitment equilibrium where gradient-based learners converge to a suboptimal cooperation level. The Nash Trap is empirically universal across five algorithm families, five architecture families, 400+ hyperparameter configurations, and 12 multiplier conditions.</p><p><strong>Theory (proved):</strong> Signal dilution and basin dominance establish Price of Anarchy PoA ≥ Ω(N) -- polynomial divergence contrasting sharply with O(1) in standard social dilemmas.</p><p><strong>Theory (conjectured):</strong> Escape may require Ω(e^{cN}) gradient evaluations under Freidlin-Wentzell heuristics.</p><p><strong>Prescription:</strong> Commitment floors bypass the trap in O(1) steps. MACCL learns state-dependent floors adaptively (+241% welfare in Cleanup).</p><p><strong>Validation:</strong> Coin Game augmented with TPSD resource collapse (160 seeds, 22%→78% survival), Gordon-Schaefer fishery with Allee-effect tipping point (300 seeds, 17%→88%), and DeepMind's Melting Pot clean_up (n=80, d=0.26, boundary-consistent directional evidence, p=0.10 two-tailed).</p><p><strong>Change log (v6 vs v5):</strong> Melting Pot expansion completed to n=80 (from n=25); PoA proof sketch restructured (bound a/b separation, alpha quantified); theorem/proposition naming reconciled; 5 rounds of independent cold review incorporated.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19649220 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | Commitment Floors for Tipping-Point Commons: Escaping Nash Traps in Multi-Agent Reinforcement Learning Author, Anonymous Multi-agent reinforcement learning MARL Nash equilibrium Social dilemma Commitment device Tipping point Price of Anarchy Cooperation [Note 2026-05-18: This record's attached file name contains legacy internal naming that includes a venue label; the file name does not represent any active submission status, and the attached content is an anonymous preprint. File renaming is not permitted on published Zenodo records; a Zenodo Support request for either file rename or record withdrawal has been submitted.]<br><br><p><strong>Status:</strong> [redacted venue] [redacted venue] under double-blind review. Author identity anonymized.</p><p>We characterize cooperation failure in <strong>Tipping-Point Social Dilemmas</strong> (TPSDs) -- multi-agent settings where thresholded resource dynamics can trigger irreversible collapse. We identify the <strong>Nash Trap</strong>: a locally stable under-commitment equilibrium where gradient-based learners converge to a suboptimal cooperation level. The Nash Trap is empirically universal across five algorithm families, five architecture families, 400+ hyperparameter configurations, and 12 multiplier conditions.</p><p><strong>Theory (proved):</strong> Signal dilution and basin dominance establish Price of Anarchy PoA ≥ Ω(N) -- polynomial divergence contrasting sharply with O(1) in standard social dilemmas.</p><p><strong>Theory (conjectured):</strong> Escape may require Ω(e^{cN}) gradient evaluations under Freidlin-Wentzell heuristics.</p><p><strong>Prescription:</strong> Commitment floors bypass the trap in O(1) steps. MACCL learns state-dependent floors adaptively (+241% welfare in Cleanup).</p><p><strong>Validation:</strong> Coin Game augmented with TPSD resource collapse (160 seeds, 22%→78% survival), Gordon-Schaefer fishery with Allee-effect tipping point (300 seeds, 17%→88%), and DeepMind's Melting Pot clean_up (n=80, d=0.26, boundary-consistent directional evidence, p=0.10 two-tailed).</p><p><strong>Change log (v6 vs v5):</strong> Melting Pot expansion completed to n=80 (from n=25); PoA proof sketch restructured (bound a/b separation, alpha quantified); theorem/proposition naming reconciled; 5 rounds of independent cold review incorporated.</p> |
| title | Commitment Floors for Tipping-Point Commons: Escaping Nash Traps in Multi-Agent Reinforcement Learning |
| topic | Multi-agent reinforcement learning MARL Nash equilibrium Social dilemma Commitment device Tipping point Price of Anarchy Cooperation |
| url | https://doi.org/10.5281/zenodo.19649220 |