Improving Robustness of AlphaZero Algorithms to Test-Time Environment Changes
Fuente:
arXiv
Saved in:
| Main Authors: | Tamassia, Isidoro, Böhmer, Wendelin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlphaZero-Edu: Democratizing Access to AlphaZero
by: Li, Ruitong, et al.
Published: (2025)
by: Li, Ruitong, et al.
Published: (2025)
TransZero: Parallel Tree Expansion in MuZero using Transformer Networks
by: Malmsten, Emil, et al.
Published: (2025)
by: Malmsten, Emil, et al.
Published: (2025)
Diversifying AI: Towards Creative Chess with AlphaZero
by: Zahavy, Tom, et al.
Published: (2023)
by: Zahavy, Tom, et al.
Published: (2023)
Regret-Guided Search Control for Efficient Learning in AlphaZero
by: Tsai, Yun-Jui, et al.
Published: (2026)
by: Tsai, Yun-Jui, et al.
Published: (2026)
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
by: Wu, Ti-Rong, et al.
Published: (2023)
by: Wu, Ti-Rong, et al.
Published: (2023)
MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Games
by: Li, Qian-Rong, et al.
Published: (2026)
by: Li, Qian-Rong, et al.
Published: (2026)
Finding Increasingly Large Extremal Graphs with AlphaZero and Tabu Search
by: Mehrabian, Abbas, et al.
Published: (2023)
by: Mehrabian, Abbas, et al.
Published: (2023)
Towards Faster Matrix Diagonalization with Graph Isomorphism Networks and the AlphaZero Framework
by: Zollicoffer, Geigh, et al.
Published: (2024)
by: Zollicoffer, Geigh, et al.
Published: (2024)
How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2025)
by: Weltevrede, Max, et al.
Published: (2025)
Value Improved Actor Critic Algorithms
by: Oren, Yaniv, et al.
Published: (2024)
by: Oren, Yaniv, et al.
Published: (2024)
Search-contempt: a hybrid MCTS algorithm for training AlphaZero-like engines with better computational efficiency
by: Joshi, Ameya
Published: (2025)
by: Joshi, Ameya
Published: (2025)
EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control
by: Evers, Thomas, et al.
Published: (2026)
by: Evers, Thomas, et al.
Published: (2026)
Diverse Projection Ensembles for Distributional Reinforcement Learning
by: Zanger, Moritz A., et al.
Published: (2023)
by: Zanger, Moritz A., et al.
Published: (2023)
Epistemic Monte Carlo Tree Search
by: Oren, Yaniv, et al.
Published: (2022)
by: Oren, Yaniv, et al.
Published: (2022)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
Explore-Go: Leveraging Exploration for Generalisation in Deep Reinforcement Learning
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
Generalisation to unseen topologies: Towards control of biological neural network activity
by: Engwegen, Laurens, et al.
Published: (2024)
by: Engwegen, Laurens, et al.
Published: (2024)
Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver
by: Sherwood, Joshua, et al.
Published: (2026)
by: Sherwood, Joshua, et al.
Published: (2026)
Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model
by: Zanger, Moritz A., et al.
Published: (2025)
by: Zanger, Moritz A., et al.
Published: (2025)
Twice Sequential Monte Carlo for Tree Search
by: Oren, Yaniv, et al.
Published: (2025)
by: Oren, Yaniv, et al.
Published: (2025)
On the Equivalence of Random Network Distillation, Deep Ensembles, and Bayesian Inference
by: Zanger, Moritz A., et al.
Published: (2026)
by: Zanger, Moritz A., et al.
Published: (2026)
Representation Matters for Mastering Chess: Improved Feature Representation in AlphaZero Outperforms Switching to Transformers
by: Czech, Johannes, et al.
Published: (2023)
by: Czech, Johannes, et al.
Published: (2023)
ShortCircuit: AlphaZero-Driven Circuit Design
by: Tsaras, Dimitrios, et al.
Published: (2024)
by: Tsaras, Dimitrios, et al.
Published: (2024)
Universal Value-Function Uncertainties
by: Zanger, Moritz A., et al.
Published: (2025)
by: Zanger, Moritz A., et al.
Published: (2025)
Change of Thought: Adaptive Test-Time Computation
by: Mathur, Mrinal, et al.
Published: (2025)
by: Mathur, Mrinal, et al.
Published: (2025)
Curricula for Learning Robust Policies with Factored State Representations in Changing Environments
by: Panayiotou, Panayiotis, et al.
Published: (2024)
by: Panayiotou, Panayiotis, et al.
Published: (2024)
A Penalty-Based Guardrail Algorithm for Non-Decreasing Optimization with Inequality Constraints
by: Stepanovic, Ksenija, et al.
Published: (2024)
by: Stepanovic, Ksenija, et al.
Published: (2024)
Reproducing AlphaZero on Tablut: Self-Play RL for an Asymmetric Board Game
by: Lees, Tõnis, et al.
Published: (2026)
by: Lees, Tõnis, et al.
Published: (2026)
Active Test-Time Adaptation: Theoretical Analyses and An Algorithm
by: Gui, Shurui, et al.
Published: (2024)
by: Gui, Shurui, et al.
Published: (2024)
You Shall Pass: Dealing with the Zero-Gradient Problem in Predict and Optimize for Convex Optimization
by: Veviurko, Grigorii, et al.
Published: (2023)
by: Veviurko, Grigorii, et al.
Published: (2023)
TNT: Improving Chunkwise Training for Test-Time Memorization
by: Li, Zeman, et al.
Published: (2025)
by: Li, Zeman, et al.
Published: (2025)
The Over-Certainty Phenomenon in Modern Test-Time Adaptation Algorithms
by: Amin, Fin, et al.
Published: (2024)
by: Amin, Fin, et al.
Published: (2024)
Decocted Experience Improves Test-Time Inference in LLM Agents
by: Shen, Maohao, et al.
Published: (2026)
by: Shen, Maohao, et al.
Published: (2026)
Computer Architecture's AlphaZero Moment: Automated Discovery in an Encircled World
by: Sankaralingam, Karthikeyan
Published: (2026)
by: Sankaralingam, Karthikeyan
Published: (2026)
Zero-Overhead Introspection for Adaptive Test-Time Compute
by: Manvi, Rohin, et al.
Published: (2025)
by: Manvi, Rohin, et al.
Published: (2025)
EVIL: Evolving Interpretable Algorithms for Zero-Shot Inference on Event Sequences and Time Series with LLMs
by: Berghaus, David
Published: (2026)
by: Berghaus, David
Published: (2026)
MantisV2: Closing the Zero-Shot Gap in Time Series Classification with Synthetic Data and Test-Time Strategies
by: Feofanov, Vasilii, et al.
Published: (2026)
by: Feofanov, Vasilii, et al.
Published: (2026)
Modular Recurrence in Contextual MDPs for Universal Morphology Control
by: Engwegen, Laurens, et al.
Published: (2025)
by: Engwegen, Laurens, et al.
Published: (2025)
Test-Time Training for Zero-Resource Dense Retrieval Reranking
by: Liu, Shiyan, et al.
Published: (2026)
by: Liu, Shiyan, et al.
Published: (2026)
Enabling Self-Improving Agents to Learn at Test Time With Human-In-The-Loop Guidance
by: He, Yufei, et al.
Published: (2025)
by: He, Yufei, et al.
Published: (2025)
Similar Items
-
AlphaZero-Edu: Democratizing Access to AlphaZero
by: Li, Ruitong, et al.
Published: (2025) -
TransZero: Parallel Tree Expansion in MuZero using Transformer Networks
by: Malmsten, Emil, et al.
Published: (2025) -
Diversifying AI: Towards Creative Chess with AlphaZero
by: Zahavy, Tom, et al.
Published: (2023) -
Regret-Guided Search Control for Efficient Learning in AlphaZero
by: Tsai, Yun-Jui, et al.
Published: (2026) -
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
by: Wu, Ti-Rong, et al.
Published: (2023)