MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Games
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Qian-Rong, Guei, Hung, Wu, I-Chen, Wu, Ti-Rong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
by: Wu, Ti-Rong, et al.
Published: (2023)
by: Wu, Ti-Rong, et al.
Published: (2023)
Regret-Guided Search Control for Efficient Learning in AlphaZero
by: Tsai, Yun-Jui, et al.
Published: (2026)
by: Tsai, Yun-Jui, et al.
Published: (2026)
AlphaZero-Edu: Democratizing Access to AlphaZero
by: Li, Ruitong, et al.
Published: (2025)
by: Li, Ruitong, et al.
Published: (2025)
Demystifying MuZero Planning: Interpreting the Learned Model
by: Guei, Hung, et al.
Published: (2024)
by: Guei, Hung, et al.
Published: (2024)
OptionZero: Planning with Learned Options
by: Huang, Po-Wei, et al.
Published: (2025)
by: Huang, Po-Wei, et al.
Published: (2025)
Evaluating Game Difficulty in Tetris Block Puzzle
by: Wang, Chun-Jui, et al.
Published: (2026)
by: Wang, Chun-Jui, et al.
Published: (2026)
Diversifying AI: Towards Creative Chess with AlphaZero
by: Zahavy, Tom, et al.
Published: (2023)
by: Zahavy, Tom, et al.
Published: (2023)
Improving Robustness of AlphaZero Algorithms to Test-Time Environment Changes
by: Tamassia, Isidoro, et al.
Published: (2025)
by: Tamassia, Isidoro, et al.
Published: (2025)
A Novel Approach to Solving Goal-Achieving Problems for Board Games
by: Shih, Chung-Chin, et al.
Published: (2021)
by: Shih, Chung-Chin, et al.
Published: (2021)
Dynamic Sight Range Selection in Multi-Agent Reinforcement Learning
by: Liao, Wei-Chen, et al.
Published: (2025)
by: Liao, Wei-Chen, et al.
Published: (2025)
Finding Increasingly Large Extremal Graphs with AlphaZero and Tabu Search
by: Mehrabian, Abbas, et al.
Published: (2023)
by: Mehrabian, Abbas, et al.
Published: (2023)
Towards Faster Matrix Diagonalization with Graph Isomorphism Networks and the AlphaZero Framework
by: Zollicoffer, Geigh, et al.
Published: (2024)
by: Zollicoffer, Geigh, et al.
Published: (2024)
Bridging Local and Global Knowledge via Transformer in Board Games
by: Ju, Yan-Ru, et al.
Published: (2024)
by: Ju, Yan-Ru, et al.
Published: (2024)
Strength Estimation and Human-Like Strength Adjustment in Games
by: Chen, Chun Jung, et al.
Published: (2025)
by: Chen, Chun Jung, et al.
Published: (2025)
Search-contempt: a hybrid MCTS algorithm for training AlphaZero-like engines with better computational efficiency
by: Joshi, Ameya
Published: (2025)
by: Joshi, Ameya
Published: (2025)
Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver
by: Sherwood, Joshua, et al.
Published: (2026)
by: Sherwood, Joshua, et al.
Published: (2026)
Relevance-Zone Reduction in Game Solving
by: Lin, Chi-Huang, et al.
Published: (2025)
by: Lin, Chi-Huang, et al.
Published: (2025)
Reproducing AlphaZero on Tablut: Self-Play RL for an Asymmetric Board Game
by: Lees, Tõnis, et al.
Published: (2026)
by: Lees, Tõnis, et al.
Published: (2026)
A Study of Solving Life-and-Death Problems in Go Using Relevance-Zone Based Solvers
by: Shih, Chung-Chin, et al.
Published: (2025)
by: Shih, Chung-Chin, et al.
Published: (2025)
MAPLE: Metadata Augmented Private Language Evolution
by: Chien, Eli, et al.
Published: (2026)
by: Chien, Eli, et al.
Published: (2026)
ShortCircuit: AlphaZero-Driven Circuit Design
by: Tsaras, Dimitrios, et al.
Published: (2024)
by: Tsaras, Dimitrios, et al.
Published: (2024)
A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence
by: Liu, Mingyang, et al.
Published: (2024)
by: Liu, Mingyang, et al.
Published: (2024)
Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
by: Guo, Jian-Ting, et al.
Published: (2025)
by: Guo, Jian-Ting, et al.
Published: (2025)
Perceptual Similarity for Measuring Decision-Making Style and Policy Diversity in Games
by: Lin, Chiu-Chou, et al.
Published: (2024)
by: Lin, Chiu-Chou, et al.
Published: (2024)
AlphaZero Neural Scaling and Zipf's Law: a Tale of Board Games and Power Laws
by: Neumann, Oren, et al.
Published: (2024)
by: Neumann, Oren, et al.
Published: (2024)
Policy-regularized Offline Multi-objective Reinforcement Learning
by: Lin, Qian, et al.
Published: (2024)
by: Lin, Qian, et al.
Published: (2024)
Solving 7x7 Killall-Go with Seki Database
by: Tsai, Yun-Jui, et al.
Published: (2024)
by: Tsai, Yun-Jui, et al.
Published: (2024)
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
by: Huang, Jiawei, et al.
Published: (2025)
by: Huang, Jiawei, et al.
Published: (2025)
Learning Game-Playing Agents with Generative Code Optimization
by: Kuang, Zhiyi, et al.
Published: (2025)
by: Kuang, Zhiyi, et al.
Published: (2025)
State Space Models over Directed Graphs
by: She, Junzhi, et al.
Published: (2025)
by: She, Junzhi, et al.
Published: (2025)
Efficiently Training Neural Networks for Imperfect Information Games by Sampling Information Sets
by: Bertram, Timo, et al.
Published: (2024)
by: Bertram, Timo, et al.
Published: (2024)
Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization
by: Liu, Zongkai, et al.
Published: (2024)
by: Liu, Zongkai, et al.
Published: (2024)
PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning
by: Wu, Feijie, et al.
Published: (2025)
by: Wu, Feijie, et al.
Published: (2025)
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
by: Shang, Shuning, et al.
Published: (2026)
by: Shang, Shuning, et al.
Published: (2026)
OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
MAPLE: Micro Analysis of Pairwise Language Evolution for Few-Shot Claim Verification
by: Zeng, Xia, et al.
Published: (2024)
by: Zeng, Xia, et al.
Published: (2024)
MAPLE: A Framework for Active Preference Learning Guided by Large Language Models
by: Mahmud, Saaduddin, et al.
Published: (2024)
by: Mahmud, Saaduddin, et al.
Published: (2024)
Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational Data
by: Leung, Cheuk Hang, et al.
Published: (2025)
by: Leung, Cheuk Hang, et al.
Published: (2025)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
Joint Task Offloading, Inference Optimization and UAV Trajectory Planning for Generative AI Empowered Intelligent Transportation Digital Twin
by: Li, Xiaohuan, et al.
Published: (2026)
by: Li, Xiaohuan, et al.
Published: (2026)
Similar Items
-
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
by: Wu, Ti-Rong, et al.
Published: (2023) -
Regret-Guided Search Control for Efficient Learning in AlphaZero
by: Tsai, Yun-Jui, et al.
Published: (2026) -
AlphaZero-Edu: Democratizing Access to AlphaZero
by: Li, Ruitong, et al.
Published: (2025) -
Demystifying MuZero Planning: Interpreting the Learned Model
by: Guei, Hung, et al.
Published: (2024) -
OptionZero: Planning with Learned Options
by: Huang, Po-Wei, et al.
Published: (2025)