Reproducing AlphaZero on Tablut: Self-Play RL for an Asymmetric Board Game
Fuente:
arXiv
Saved in:
| Main Authors: | Lees, Tõnis, Matiisen, Tambet |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlphaZero-Edu: Democratizing Access to AlphaZero
by: Li, Ruitong, et al.
Published: (2025)
by: Li, Ruitong, et al.
Published: (2025)
AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws
by: Neumann, Oren, et al.
Published: (2024)
by: Neumann, Oren, et al.
Published: (2024)
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
by: Wu, Ti-Rong, et al.
Published: (2023)
by: Wu, Ti-Rong, et al.
Published: (2023)
MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Games
by: Li, Qian-Rong, et al.
Published: (2026)
by: Li, Qian-Rong, et al.
Published: (2026)
ShortCircuit: AlphaZero-Driven Circuit Design
by: Tsaras, Dimitrios, et al.
Published: (2024)
by: Tsaras, Dimitrios, et al.
Published: (2024)
Diversifying AI: Towards Creative Chess with AlphaZero
by: Zahavy, Tom, et al.
Published: (2023)
by: Zahavy, Tom, et al.
Published: (2023)
Regret-Guided Search Control for Efficient Learning in AlphaZero
by: Tsai, Yun-Jui, et al.
Published: (2026)
by: Tsai, Yun-Jui, et al.
Published: (2026)
Improving Robustness of AlphaZero Algorithms to Test-Time Environment Changes
by: Tamassia, Isidoro, et al.
Published: (2025)
by: Tamassia, Isidoro, et al.
Published: (2025)
Computer Architecture's AlphaZero Moment: Automated Discovery in an Encircled World
by: Sankaralingam, Karthikeyan
Published: (2026)
by: Sankaralingam, Karthikeyan
Published: (2026)
Finding Increasingly Large Extremal Graphs with AlphaZero and Tabu Search
by: Mehrabian, Abbas, et al.
Published: (2023)
by: Mehrabian, Abbas, et al.
Published: (2023)
Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver
by: Sherwood, Joshua, et al.
Published: (2026)
by: Sherwood, Joshua, et al.
Published: (2026)
Towards Faster Matrix Diagonalization with Graph Isomorphism Networks and the AlphaZero Framework
by: Zollicoffer, Geigh, et al.
Published: (2024)
by: Zollicoffer, Geigh, et al.
Published: (2024)
Deep Hedging Under Non-Convexity: Limitations and a Case for AlphaZero
by: Maggiolo, Matteo, et al.
Published: (2025)
by: Maggiolo, Matteo, et al.
Published: (2025)
Search-contempt: a hybrid MCTS algorithm for training AlphaZero-like engines with better computational efficiency
by: Joshi, Ameya
Published: (2025)
by: Joshi, Ameya
Published: (2025)
Playing Board Games with the Predict Results of Beam Search Algorithm
by: Pastukhov, Sergey
Published: (2024)
by: Pastukhov, Sergey
Published: (2024)
Simultaneous AlphaZero: Extending Tree Search to Markov Games
by: Becker, Tyler, et al.
Published: (2025)
by: Becker, Tyler, et al.
Published: (2025)
GASP: Guided Asymmetric Self-Play For Coding LLMs
by: Jana, Swadesh, et al.
Published: (2026)
by: Jana, Swadesh, et al.
Published: (2026)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
by: Tan, Zhewen, et al.
Published: (2026)
by: Tan, Zhewen, et al.
Published: (2026)
Learning to Drive via Asymmetric Self-Play
by: Zhang, Chris, et al.
Published: (2024)
by: Zhang, Chris, et al.
Published: (2024)
Internalizing World Models via Self-Play Finetuning for Agentic RL
by: Chen, Shiqi, et al.
Published: (2025)
by: Chen, Shiqi, et al.
Published: (2025)
Unitary Synthesis with AlphaZero via Dynamic Circuits
by: Valcarce, Xavier, et al.
Published: (2025)
by: Valcarce, Xavier, et al.
Published: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Explore to Generalize in Zero-Shot RL
by: Zisselman, Ev, et al.
Published: (2023)
by: Zisselman, Ev, et al.
Published: (2023)
Offline Fictitious Self-Play for Competitive Games
by: Chen, Jingxiao, et al.
Published: (2024)
by: Chen, Jingxiao, et al.
Published: (2024)
AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play
by: Murgoci, Vlad, et al.
Published: (2026)
by: Murgoci, Vlad, et al.
Published: (2026)
Fast and Furious Symmetric Learning in Zero-Sum Games: Gradient Descent as Fictitious Play
by: Lazarsfeld, John, et al.
Published: (2025)
by: Lazarsfeld, John, et al.
Published: (2025)
G-Zero: Self-Play for Open-Ended Generation from Zero Data
by: Huang, Chengsong, et al.
Published: (2026)
by: Huang, Chengsong, et al.
Published: (2026)
Scaling Self-Play with Self-Guidance
by: Bailey, Luke, et al.
Published: (2026)
by: Bailey, Luke, et al.
Published: (2026)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Agents Play Thousands of 3D Video Games
by: Xu, Zhongwen, et al.
Published: (2025)
by: Xu, Zhongwen, et al.
Published: (2025)
Boardwalk: Towards a Framework for Creating Board Games with LLMs
by: Becker, Álvaro Guglielmin, et al.
Published: (2025)
by: Becker, Álvaro Guglielmin, et al.
Published: (2025)
Mastering NIM and Impartial Games with Weak Neural Networks: An AlphaZero-inspired Multi-Frame Approach
by: Riis, Søren
Published: (2024)
by: Riis, Søren
Published: (2024)
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
by: Chen, Jiaqi, et al.
Published: (2025)
by: Chen, Jiaqi, et al.
Published: (2025)
Read to Play (R2-Play): Decision Transformer with Multimodal Game Instruction
by: Jin, Yonggang, et al.
Published: (2024)
by: Jin, Yonggang, et al.
Published: (2024)
Playing Non-Embedded Card-Based Games with Reinforcement Learning
by: Wu, Tianyang, et al.
Published: (2025)
by: Wu, Tianyang, et al.
Published: (2025)
Do We Need Transformers to Play FPS Video Games?
by: Batth, Karmanbir, et al.
Published: (2025)
by: Batth, Karmanbir, et al.
Published: (2025)
Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games
by: Lanier, JB, et al.
Published: (2026)
by: Lanier, JB, et al.
Published: (2026)
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning
by: Li, Shangzhe, et al.
Published: (2026)
by: Li, Shangzhe, et al.
Published: (2026)
Learning Game-Playing Agents with Generative Code Optimization
by: Kuang, Zhiyi, et al.
Published: (2025)
by: Kuang, Zhiyi, et al.
Published: (2025)
Bridging Local and Global Knowledge via Transformer in Board Games
by: Ju, Yan-Ru, et al.
Published: (2024)
by: Ju, Yan-Ru, et al.
Published: (2024)
Similar Items
-
AlphaZero-Edu: Democratizing Access to AlphaZero
by: Li, Ruitong, et al.
Published: (2025) -
AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws
by: Neumann, Oren, et al.
Published: (2024) -
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
by: Wu, Ti-Rong, et al.
Published: (2023) -
MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Games
by: Li, Qian-Rong, et al.
Published: (2026) -
ShortCircuit: AlphaZero-Driven Circuit Design
by: Tsaras, Dimitrios, et al.
Published: (2024)