AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Murgoci, Vlad, Spaan, Matthijs, Oren, Yaniv |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Epistemic Monte Carlo Tree Search
von: Oren, Yaniv, et al.
Veröffentlicht: (2022)
von: Oren, Yaniv, et al.
Veröffentlicht: (2022)
Trust-Region Twisted Policy Improvement
von: de Vries, Joery A., et al.
Veröffentlicht: (2025)
von: de Vries, Joery A., et al.
Veröffentlicht: (2025)
RecBayes: Recurrent Bayesian Ad Hoc Teamwork in Large Partially Observable Domains
von: Ribeiro, João G., et al.
Veröffentlicht: (2025)
von: Ribeiro, João G., et al.
Veröffentlicht: (2025)
Explore-Go: Leveraging Exploration for Generalisation in Deep Reinforcement Learning
von: Weltevrede, Max, et al.
Veröffentlicht: (2024)
von: Weltevrede, Max, et al.
Veröffentlicht: (2024)
PMCTS: Particle Monte Carlo Tree Search for Principled Parallelized Inference Time Scaling
von: Oren, Yaniv, et al.
Veröffentlicht: (2026)
von: Oren, Yaniv, et al.
Veröffentlicht: (2026)
Twice Sequential Monte Carlo for Tree Search
von: Oren, Yaniv, et al.
Veröffentlicht: (2025)
von: Oren, Yaniv, et al.
Veröffentlicht: (2025)
VariBASed: Variational Bayes-Adaptive Sequential Monte-Carlo Planning for Deep Reinforcement Learning
von: de Vries, Joery A., et al.
Veröffentlicht: (2026)
von: de Vries, Joery A., et al.
Veröffentlicht: (2026)
Value Improved Actor Critic Algorithms
von: Oren, Yaniv, et al.
Veröffentlicht: (2024)
von: Oren, Yaniv, et al.
Veröffentlicht: (2024)
Universal Value-Function Uncertainties
von: Zanger, Moritz A., et al.
Veröffentlicht: (2025)
von: Zanger, Moritz A., et al.
Veröffentlicht: (2025)
Diverse Projection Ensembles for Distributional Reinforcement Learning
von: Zanger, Moritz A., et al.
Veröffentlicht: (2023)
von: Zanger, Moritz A., et al.
Veröffentlicht: (2023)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
von: Li, Guopeng, et al.
Veröffentlicht: (2026)
von: Li, Guopeng, et al.
Veröffentlicht: (2026)
Bad Habits: Policy Confounding and Out-of-Trajectory Generalization in RL
von: Suau, Miguel, et al.
Veröffentlicht: (2023)
von: Suau, Miguel, et al.
Veröffentlicht: (2023)
Positive Experience Reflection for Agents in Interactive Text Environments
von: Lippmann, Philip, et al.
Veröffentlicht: (2024)
von: Lippmann, Philip, et al.
Veröffentlicht: (2024)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
von: Weltevrede, Max, et al.
Veröffentlicht: (2025)
von: Weltevrede, Max, et al.
Veröffentlicht: (2025)
Priors Matter: Addressing Misspecification in Bayesian Deep Q-Learning
von: van der Vaart, Pascal R., et al.
Veröffentlicht: (2025)
von: van der Vaart, Pascal R., et al.
Veröffentlicht: (2025)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
von: Weltevrede, Max, et al.
Veröffentlicht: (2024)
von: Weltevrede, Max, et al.
Veröffentlicht: (2024)
Bayesian Meta-Reinforcement Learning with Laplace Variational Recurrent Networks
von: de Vries, Joery A., et al.
Veröffentlicht: (2025)
von: de Vries, Joery A., et al.
Veröffentlicht: (2025)
Sparse Masked Attention Policies for Reliable Generalization
von: Horsch, Caroline, et al.
Veröffentlicht: (2026)
von: Horsch, Caroline, et al.
Veröffentlicht: (2026)
SpinGPT: A Large-Language-Model Approach to Playing Poker Correctly
von: Maugin, Narada, et al.
Veröffentlicht: (2025)
von: Maugin, Narada, et al.
Veröffentlicht: (2025)
Reinforcement Learning by Guided Safe Exploration
von: Yang, Qisong, et al.
Veröffentlicht: (2023)
von: Yang, Qisong, et al.
Veröffentlicht: (2023)
No-Regret Learning of Nash Equilibrium for Black-Box Games via Gaussian Processes
von: Han, Minbiao, et al.
Veröffentlicht: (2024)
von: Han, Minbiao, et al.
Veröffentlicht: (2024)
When Do Off-Policy and On-Policy Policy Gradient Methods Align?
von: Mambelli, Davide, et al.
Veröffentlicht: (2024)
von: Mambelli, Davide, et al.
Veröffentlicht: (2024)
Graph Learning Is Suboptimal in Causal Bandits
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2025)
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2025)
Safe Multi-Agent Reinforcement Learning with Convergence to Generalized Nash Equilibrium
von: Li, Zeyang, et al.
Veröffentlicht: (2024)
von: Li, Zeyang, et al.
Veröffentlicht: (2024)
AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws
von: Neumann, Oren, et al.
Veröffentlicht: (2024)
von: Neumann, Oren, et al.
Veröffentlicht: (2024)
Accelerating Nash Equilibrium Convergence in Monte Carlo Settings Through Counterfactual Value Based Fictitious Play
von: Qi, Ju, et al.
Veröffentlicht: (2023)
von: Qi, Ju, et al.
Veröffentlicht: (2023)
Generalizing Beyond Suboptimality: Offline Reinforcement Learning Learns Effective Scheduling through Random Data
von: van Remmerden, Jesse, et al.
Veröffentlicht: (2025)
von: van Remmerden, Jesse, et al.
Veröffentlicht: (2025)
An Online Feasible Point Method for Benign Generalized Nash Equilibrium Problems
von: Sachs, Sarah, et al.
Veröffentlicht: (2024)
von: Sachs, Sarah, et al.
Veröffentlicht: (2024)
SORREL: Suboptimal-Demonstration-Guided Reinforcement Learning for Learning to Branch
von: Feng, Shengyu, et al.
Veröffentlicht: (2024)
von: Feng, Shengyu, et al.
Veröffentlicht: (2024)
In-Context Reinforcement Learning From Suboptimal Historical Data
von: Dong, Juncheng, et al.
Veröffentlicht: (2026)
von: Dong, Juncheng, et al.
Veröffentlicht: (2026)
Mixed Strategy Nash Equilibrium for Crowd Navigation
von: Sun, Max Muchen, et al.
Veröffentlicht: (2024)
von: Sun, Max Muchen, et al.
Veröffentlicht: (2024)
Super-Exponential Regret for UCT, AlphaGo and Variants
von: Orseau, Laurent, et al.
Veröffentlicht: (2024)
von: Orseau, Laurent, et al.
Veröffentlicht: (2024)
Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model
von: Zanger, Moritz A., et al.
Veröffentlicht: (2025)
von: Zanger, Moritz A., et al.
Veröffentlicht: (2025)
Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked Systems
von: Suau, Miguel, et al.
Veröffentlicht: (2022)
von: Suau, Miguel, et al.
Veröffentlicht: (2022)
From Debate to Equilibrium: Belief-Driven Multi-Agent LLM Reasoning via Bayesian Nash Equilibrium
von: Yi, Xie, et al.
Veröffentlicht: (2025)
von: Yi, Xie, et al.
Veröffentlicht: (2025)
Equilibrium Selection in Multi-Agent Policy Gradients via Opponent-Aware Basin Entry
von: Shcherbinin, Yevhen, et al.
Veröffentlicht: (2026)
von: Shcherbinin, Yevhen, et al.
Veröffentlicht: (2026)
Towards General Preference Alignment: Diffusion Models at Nash Equilibrium
von: Hu, Jiaming, et al.
Veröffentlicht: (2026)
von: Hu, Jiaming, et al.
Veröffentlicht: (2026)
Convergence to Nash Equilibrium and No-regret Guarantee in (Markov) Potential Games
von: Dong, Jing, et al.
Veröffentlicht: (2024)
von: Dong, Jing, et al.
Veröffentlicht: (2024)
Large-Scale Auto-bidding with Nash Equilibrium Constraints
von: Mou, Zhiyu, et al.
Veröffentlicht: (2025)
von: Mou, Zhiyu, et al.
Veröffentlicht: (2025)
Reproducing AlphaZero on Tablut: Self-Play RL for an Asymmetric Board Game
von: Lees, Tõnis, et al.
Veröffentlicht: (2026)
von: Lees, Tõnis, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Epistemic Monte Carlo Tree Search
von: Oren, Yaniv, et al.
Veröffentlicht: (2022) -
Trust-Region Twisted Policy Improvement
von: de Vries, Joery A., et al.
Veröffentlicht: (2025) -
RecBayes: Recurrent Bayesian Ad Hoc Teamwork in Large Partially Observable Domains
von: Ribeiro, João G., et al.
Veröffentlicht: (2025) -
Explore-Go: Leveraging Exploration for Generalisation in Deep Reinforcement Learning
von: Weltevrede, Max, et al.
Veröffentlicht: (2024) -
PMCTS: Particle Monte Carlo Tree Search for Principled Parallelized Inference Time Scaling
von: Oren, Yaniv, et al.
Veröffentlicht: (2026)