Regret-Guided Search Control for Efficient Learning in AlphaZero
Fuente:
arXiv
Salvato in:
| Autori principali: | Tsai, Yun-Jui, Chen, Wei-Yu, Ju, Yan-Ru, Chang, Yu-Hung, Wu, Ti-Rong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
di: Wu, Ti-Rong, et al.
Pubblicazione: (2023)
di: Wu, Ti-Rong, et al.
Pubblicazione: (2023)
MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Games
di: Li, Qian-Rong, et al.
Pubblicazione: (2026)
di: Li, Qian-Rong, et al.
Pubblicazione: (2026)
AlphaZero-Edu: Democratizing Access to AlphaZero
di: Li, Ruitong, et al.
Pubblicazione: (2025)
di: Li, Ruitong, et al.
Pubblicazione: (2025)
Demystifying MuZero Planning: Interpreting the Learned Model
di: Guei, Hung, et al.
Pubblicazione: (2024)
di: Guei, Hung, et al.
Pubblicazione: (2024)
Finding Increasingly Large Extremal Graphs with AlphaZero and Tabu Search
di: Mehrabian, Abbas, et al.
Pubblicazione: (2023)
di: Mehrabian, Abbas, et al.
Pubblicazione: (2023)
Diversifying AI: Towards Creative Chess with AlphaZero
di: Zahavy, Tom, et al.
Pubblicazione: (2023)
di: Zahavy, Tom, et al.
Pubblicazione: (2023)
Improving Robustness of AlphaZero Algorithms to Test-Time Environment Changes
di: Tamassia, Isidoro, et al.
Pubblicazione: (2025)
di: Tamassia, Isidoro, et al.
Pubblicazione: (2025)
Search-contempt: a hybrid MCTS algorithm for training AlphaZero-like engines with better computational efficiency
di: Joshi, Ameya
Pubblicazione: (2025)
di: Joshi, Ameya
Pubblicazione: (2025)
OptionZero: Planning with Learned Options
di: Huang, Po-Wei, et al.
Pubblicazione: (2025)
di: Huang, Po-Wei, et al.
Pubblicazione: (2025)
Relevance-Zone Reduction in Game Solving
di: Lin, Chi-Huang, et al.
Pubblicazione: (2025)
di: Lin, Chi-Huang, et al.
Pubblicazione: (2025)
Representation Matters for Mastering Chess: Improved Feature Representation in AlphaZero Outperforms Switching to Transformers
di: Czech, Johannes, et al.
Pubblicazione: (2023)
di: Czech, Johannes, et al.
Pubblicazione: (2023)
Solving 7x7 Killall-Go with Seki Database
di: Tsai, Yun-Jui, et al.
Pubblicazione: (2024)
di: Tsai, Yun-Jui, et al.
Pubblicazione: (2024)
Towards Faster Matrix Diagonalization with Graph Isomorphism Networks and the AlphaZero Framework
di: Zollicoffer, Geigh, et al.
Pubblicazione: (2024)
di: Zollicoffer, Geigh, et al.
Pubblicazione: (2024)
Bridging Local and Global Knowledge via Transformer in Board Games
di: Ju, Yan-Ru, et al.
Pubblicazione: (2024)
di: Ju, Yan-Ru, et al.
Pubblicazione: (2024)
TSS GAZ PTP: Towards Improving Gumbel AlphaZero with Two-stage Self-play for Multi-constrained Electric Vehicle Routing Problems
di: Wang, Hui, et al.
Pubblicazione: (2025)
di: Wang, Hui, et al.
Pubblicazione: (2025)
Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver
di: Sherwood, Joshua, et al.
Pubblicazione: (2026)
di: Sherwood, Joshua, et al.
Pubblicazione: (2026)
A Study of Solving Life-and-Death Problems in Go Using Relevance-Zone Based Solvers
di: Shih, Chung-Chin, et al.
Pubblicazione: (2025)
di: Shih, Chung-Chin, et al.
Pubblicazione: (2025)
Evaluating Game Difficulty in Tetris Block Puzzle
di: Wang, Chun-Jui, et al.
Pubblicazione: (2026)
di: Wang, Chun-Jui, et al.
Pubblicazione: (2026)
Dynamic Sight Range Selection in Multi-Agent Reinforcement Learning
di: Liao, Wei-Chen, et al.
Pubblicazione: (2025)
di: Liao, Wei-Chen, et al.
Pubblicazione: (2025)
Annealing Self-Distillation Rectification Improves Adversarial Training
di: Wu, Yu-Yu, et al.
Pubblicazione: (2023)
di: Wu, Yu-Yu, et al.
Pubblicazione: (2023)
Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided Search
di: Liu, Max, et al.
Pubblicazione: (2024)
di: Liu, Max, et al.
Pubblicazione: (2024)
Alpha Discovery via Grammar-Guided Learning and Search
di: Yang, Han, et al.
Pubblicazione: (2026)
di: Yang, Han, et al.
Pubblicazione: (2026)
Mastering NIM and Impartial Games with Weak Neural Networks: An AlphaZero-inspired Multi-Frame Approach
di: Riis, Søren
Pubblicazione: (2024)
di: Riis, Søren
Pubblicazione: (2024)
Simultaneous AlphaZero: Extending Tree Search to Markov Games
di: Becker, Tyler, et al.
Pubblicazione: (2025)
di: Becker, Tyler, et al.
Pubblicazione: (2025)
Super-Exponential Regret for UCT, AlphaGo and Variants
di: Orseau, Laurent, et al.
Pubblicazione: (2024)
di: Orseau, Laurent, et al.
Pubblicazione: (2024)
LITA: An Efficient LLM-assisted Iterative Topic Augmentation Framework
di: Chang, Chia-Hsuan, et al.
Pubblicazione: (2024)
di: Chang, Chia-Hsuan, et al.
Pubblicazione: (2024)
CFDA & CLIP at TREC iKAT 2025: Enhancing Personalized Conversational Search via Query Reformulation and Rank Fusion
di: Chang, Yu-Cheng, et al.
Pubblicazione: (2025)
di: Chang, Yu-Cheng, et al.
Pubblicazione: (2025)
PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting
di: Tsai, Yu-Che, et al.
Pubblicazione: (2026)
di: Tsai, Yu-Che, et al.
Pubblicazione: (2026)
HNote: Extending YNote with Hexadecimal Encoding for Fine-Tuning LLMs in Music Modeling
di: Chu, Hung-Ying, et al.
Pubblicazione: (2025)
di: Chu, Hung-Ying, et al.
Pubblicazione: (2025)
Inteligencia Artificial y Juegos de Tablero: Desde el Turco hasta AlphaZero
di: Ivan Francisco Valencia
Pubblicazione: (2022)
di: Ivan Francisco Valencia
Pubblicazione: (2022)
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
di: Tan, Jiejun, et al.
Pubblicazione: (2025)
di: Tan, Jiejun, et al.
Pubblicazione: (2025)
A Novel Approach to Solving Goal-Achieving Problems for Board Games
di: Shih, Chung-Chin, et al.
Pubblicazione: (2021)
di: Shih, Chung-Chin, et al.
Pubblicazione: (2021)
Quantum-Enhanced Parameter-Efficient Learning for Typhoon Trajectory Forecasting
di: Liu, Chen-Yu, et al.
Pubblicazione: (2025)
di: Liu, Chen-Yu, et al.
Pubblicazione: (2025)
RailEstate: An Interactive System for Metro Linked Property Trends
di: Chang, Chen-Wei, et al.
Pubblicazione: (2025)
di: Chang, Chen-Wei, et al.
Pubblicazione: (2025)
ShortCircuit: AlphaZero-Driven Circuit Design
di: Tsaras, Dimitrios, et al.
Pubblicazione: (2024)
di: Tsaras, Dimitrios, et al.
Pubblicazione: (2024)
Conversational Orientation Reasoning: Egocentric-to-Allocentric Navigation with Multimodal Chain-of-Thought
di: Huang, Yu Ti
Pubblicazione: (2025)
di: Huang, Yu Ti
Pubblicazione: (2025)
Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
di: Guo, Jian-Ting, et al.
Pubblicazione: (2025)
di: Guo, Jian-Ting, et al.
Pubblicazione: (2025)
Self-Attention with State-Object Weighted Combination for Compositional Zero Shot Learning
di: Chang, Cheng-Hong, et al.
Pubblicazione: (2025)
di: Chang, Cheng-Hong, et al.
Pubblicazione: (2025)
Normality-Guided Distributional Reinforcement Learning for Continuous Control
di: Byun, Ju-Seung, et al.
Pubblicazione: (2022)
di: Byun, Ju-Seung, et al.
Pubblicazione: (2022)
BTS: Bifold Teacher-Student in Semi-Supervised Learning for Indoor Two-Room Presence Detection Under Time-Varying CSI
di: Shen, Li-Hsiang, et al.
Pubblicazione: (2022)
di: Shen, Li-Hsiang, et al.
Pubblicazione: (2022)
Documenti analoghi
-
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
di: Wu, Ti-Rong, et al.
Pubblicazione: (2023) -
MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Games
di: Li, Qian-Rong, et al.
Pubblicazione: (2026) -
AlphaZero-Edu: Democratizing Access to AlphaZero
di: Li, Ruitong, et al.
Pubblicazione: (2025) -
Demystifying MuZero Planning: Interpreting the Learned Model
di: Guei, Hung, et al.
Pubblicazione: (2024) -
Finding Increasingly Large Extremal Graphs with AlphaZero and Tabu Search
di: Mehrabian, Abbas, et al.
Pubblicazione: (2023)