Equilibrium Selection in Multi-Agent Policy Gradients via Opponent-Aware Basin Entry
Fuente:
arXiv
Saved in:
| Main Authors: | Shcherbinin, Yevhen, Redina, Arina, Kalpin, Maxim, Kochetov, Vlad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Opponent Shaping in LLM Agents
by: Segura, Marta Emili Garcia, et al.
Published: (2025)
by: Segura, Marta Emili Garcia, et al.
Published: (2025)
Adaptive Opponent Policy Detection in Multi-Agent MDPs: Real-Time Strategy Switch Identification Using Running Error Estimation
by: Mridul, Mohidul Haque, et al.
Published: (2024)
by: Mridul, Mohidul Haque, et al.
Published: (2024)
Towards Sustainable Investment Policies Informed by Opponent Shaping
by: Duque, Juan Agustin, et al.
Published: (2026)
by: Duque, Juan Agustin, et al.
Published: (2026)
Optimistic Multi-Agent Policy Gradient
by: Zhao, Wenshuai, et al.
Published: (2023)
by: Zhao, Wenshuai, et al.
Published: (2023)
LOQA: Learning with Opponent Q-Learning Awareness
by: Aghajohari, Milad, et al.
Published: (2024)
by: Aghajohari, Milad, et al.
Published: (2024)
AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play
by: Murgoci, Vlad, et al.
Published: (2026)
by: Murgoci, Vlad, et al.
Published: (2026)
Differentiable Belief-based Opponent Shaping
by: Sane, Aarav G, et al.
Published: (2026)
by: Sane, Aarav G, et al.
Published: (2026)
Metric-Gradient Projection for Stable Multi-Agent Policy Learning
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
by: Reddi, Aryaman, et al.
Published: (2025)
by: Reddi, Aryaman, et al.
Published: (2025)
From Debate to Equilibrium: Belief-Driven Multi-Agent LLM Reasoning via Bayesian Nash Equilibrium
by: Yi, Xie, et al.
Published: (2025)
by: Yi, Xie, et al.
Published: (2025)
Detection of Interacting Variables for Generalized Linear Models via Neural Networks
by: Havrylenko, Yevhen, et al.
Published: (2022)
by: Havrylenko, Yevhen, et al.
Published: (2022)
A Multi-Agent, Policy-Gradient approach to Network Routing
by: Tao, Nigel, et al.
Published: (2025)
by: Tao, Nigel, et al.
Published: (2025)
Can Entry-Wise Clipping Give Spectral Control of Stochastic Gradients?
by: Song, Zitao, et al.
Published: (2026)
by: Song, Zitao, et al.
Published: (2026)
Learning to Play Against Unknown Opponents
by: Arunachaleswaran, Eshwar Ram, et al.
Published: (2024)
by: Arunachaleswaran, Eshwar Ram, et al.
Published: (2024)
Decision-making with Speculative Opponent Models
by: Sun, Jing, et al.
Published: (2022)
by: Sun, Jing, et al.
Published: (2022)
CoFi-PGMA: Counterfactual Policy Gradients under Filtered Feedback for Multi-Agent LLMs
by: Tong, Stela, et al.
Published: (2026)
by: Tong, Stela, et al.
Published: (2026)
Descent-Guided Policy Gradient for Scalable Cooperative Multi-Agent Learning
by: Yang, Shan, et al.
Published: (2026)
by: Yang, Shan, et al.
Published: (2026)
Global Convergence of Multi-Agent Policy Gradient in Markov Potential Games
by: Leonardos, Stefanos, et al.
Published: (2021)
by: Leonardos, Stefanos, et al.
Published: (2021)
EntryPrune: Neural Network Feature Selection using First Impressions
by: Zimmer, Felix, et al.
Published: (2024)
by: Zimmer, Felix, et al.
Published: (2024)
Toward Lifelong Learning in Equilibrium Propagation: Sleep-like and Awake Rehearsal for Enhanced Stability
by: Kubo, Yoshimasa, et al.
Published: (2025)
by: Kubo, Yoshimasa, et al.
Published: (2025)
PolicyEvol-Agent: Evolving Policy via Environment Perception and Self-Awareness with Theory of Mind
by: Yu, Yajie, et al.
Published: (2025)
by: Yu, Yajie, et al.
Published: (2025)
Multi-Agent Reinforcement Learning for Unmanned Aerial Vehicle Coordination by Multi-Critic Policy Gradient Optimization
by: Alon, Yoav, et al.
Published: (2020)
by: Alon, Yoav, et al.
Published: (2020)
The Phase Is the Gradient: Equilibrium Propagation for Frequency Learning in Kuramoto Networks
by: Ahmadi, Mani Rash
Published: (2026)
by: Ahmadi, Mani Rash
Published: (2026)
Cooperative Game-Theoretic Credit Assignment for Multi-Agent Policy Gradients via the Core
by: Ji, Mengda, et al.
Published: (2025)
by: Ji, Mengda, et al.
Published: (2025)
Analysing the Sample Complexity of Opponent Shaping
by: Fung, Kitty, et al.
Published: (2024)
by: Fung, Kitty, et al.
Published: (2024)
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
by: Hait, Soumita, et al.
Published: (2026)
by: Hait, Soumita, et al.
Published: (2026)
CODA: Coordination via On-Policy Diffusion for Multi-Agent Offline Reinforcement Learning
by: Hedman, Marcel, et al.
Published: (2026)
by: Hedman, Marcel, et al.
Published: (2026)
Safe Multi-Agent Reinforcement Learning with Convergence to Generalized Nash Equilibrium
by: Li, Zeyang, et al.
Published: (2024)
by: Li, Zeyang, et al.
Published: (2024)
Group Policy Gradient
by: Chen, Junhua, et al.
Published: (2025)
by: Chen, Junhua, et al.
Published: (2025)
DDEQs: Distributional Deep Equilibrium Models through Wasserstein Gradient Flows
by: Geuter, Jonathan, et al.
Published: (2025)
by: Geuter, Jonathan, et al.
Published: (2025)
GCond: Gradient Conflict Resolution via Accumulation-based Stabilization for Large-Scale Multi-Task Learning
by: Limarenko, Evgeny Alves, et al.
Published: (2025)
by: Limarenko, Evgeny Alves, et al.
Published: (2025)
Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
A Multi-Component Reward Function with Policy Gradient for Automated Feature Selection with Dynamic Regularization and Bias Mitigation
by: Khadka, Sudip, et al.
Published: (2025)
by: Khadka, Sudip, et al.
Published: (2025)
Efficient and Optimal Policy Gradient Algorithm for Corrupted Multi-armed Bandits
by: Liu, Jiayuan, et al.
Published: (2025)
by: Liu, Jiayuan, et al.
Published: (2025)
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2024)
by: Ziyin, Liu, et al.
Published: (2024)
A Basin-Selection Perspective on Grokking via Singular Learning Theory
by: Cullen, Ben, et al.
Published: (2026)
by: Cullen, Ben, et al.
Published: (2026)
Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with $k$-step Policy Gradients
by: DeWeese, Alex, et al.
Published: (2026)
by: DeWeese, Alex, et al.
Published: (2026)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
by: Corrado, Nicholas E., et al.
Published: (2023)
by: Corrado, Nicholas E., et al.
Published: (2023)
Privacy-Constrained Policies via Mutual Information Regularized Policy Gradients
by: Cundy, Chris, et al.
Published: (2020)
by: Cundy, Chris, et al.
Published: (2020)
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
by: Robert, Thomas, et al.
Published: (2024)
by: Robert, Thomas, et al.
Published: (2024)
Similar Items
-
Opponent Shaping in LLM Agents
by: Segura, Marta Emili Garcia, et al.
Published: (2025) -
Adaptive Opponent Policy Detection in Multi-Agent MDPs: Real-Time Strategy Switch Identification Using Running Error Estimation
by: Mridul, Mohidul Haque, et al.
Published: (2024) -
Towards Sustainable Investment Policies Informed by Opponent Shaping
by: Duque, Juan Agustin, et al.
Published: (2026) -
Optimistic Multi-Agent Policy Gradient
by: Zhao, Wenshuai, et al.
Published: (2023) -
LOQA: Learning with Opponent Q-Learning Awareness
by: Aghajohari, Milad, et al.
Published: (2024)