Deep Reinforcement Learning Xiangqi Player with Monte Carlo Tree Search
Fuente:
arXiv
Saved in:
| Main Authors: | Yilmaz, Berk, Hu, Junyu, Liu, Jinsong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Sparse to Dense: Toddler-inspired Reward Transition in Goal-Oriented Reinforcement Learning
by: Park, Junseok, et al.
Published: (2025)
by: Park, Junseok, et al.
Published: (2025)
Optimizing MoE Routers: Design, Implementation, and Evaluation in Transformer Models
by: Harvey, Daniel Fidel, et al.
Published: (2025)
by: Harvey, Daniel Fidel, et al.
Published: (2025)
From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning
by: Klačan, Ján, et al.
Published: (2026)
by: Klačan, Ján, et al.
Published: (2026)
Grouped Sequential Optimization Strategy -- the Application of Hyperparameter Importance Assessment in Deep Learning
by: Wang, Ruinan, et al.
Published: (2025)
by: Wang, Ruinan, et al.
Published: (2025)
Pushdown Reward Machines for Reinforcement Learning
by: Varricchione, Giovanni, et al.
Published: (2025)
by: Varricchione, Giovanni, et al.
Published: (2025)
Statistical Guarantees for Lifelong Reinforcement Learning using PAC-Bayes Theory
by: Zhang, Zhi, et al.
Published: (2024)
by: Zhang, Zhi, et al.
Published: (2024)
Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
Unsupervised Ensemble Learning Through Deep Energy-based Models
by: Maymon, Ariel, et al.
Published: (2026)
by: Maymon, Ariel, et al.
Published: (2026)
Reciprocal Learning
by: Rodemann, Julian, et al.
Published: (2024)
by: Rodemann, Julian, et al.
Published: (2024)
STACHE: Local Black-Box Explanations for Reinforcement Learning Policies
by: Elashkin, Andrew, et al.
Published: (2025)
by: Elashkin, Andrew, et al.
Published: (2025)
Large Language Model Meets Graph Neural Network in Knowledge Distillation
by: Hu, Shengxiang, et al.
Published: (2024)
by: Hu, Shengxiang, et al.
Published: (2024)
Self-Directed Task Identification
by: Gould, Timothy, et al.
Published: (2026)
by: Gould, Timothy, et al.
Published: (2026)
Convergence Dynamics and Stabilization Strategies of Co-Evolving Generative Models
by: Gao, Weiguo, et al.
Published: (2025)
by: Gao, Weiguo, et al.
Published: (2025)
Weakly Supervised Learners for Correction of AI Errors with Provable Performance Guarantees
by: Tyukin, Ivan Y., et al.
Published: (2024)
by: Tyukin, Ivan Y., et al.
Published: (2024)
Multi-State TD Target for Model-Free Reinforcement Learning
by: Wang, Wuhao, et al.
Published: (2024)
by: Wang, Wuhao, et al.
Published: (2024)
Monte Carlo Tree Search for Execution-Guided Program Repair with Large Language Models
by: Liang, Yixuan
Published: (2026)
by: Liang, Yixuan
Published: (2026)
CircuitBuilder: From Polynomials to Circuits via Reinforcement Learning
by: Zhang, Weikun K., et al.
Published: (2026)
by: Zhang, Weikun K., et al.
Published: (2026)
Topological Foundations of Reinforcement Learning
by: Kadurha, David Krame
Published: (2024)
by: Kadurha, David Krame
Published: (2024)
BatteryML:An Open-source platform for Machine Learning on Battery Degradation
by: Zhang, Han, et al.
Published: (2023)
by: Zhang, Han, et al.
Published: (2023)
A Theory of Non-Acyclic Generative Flow Networks
by: Brunswic, Leo Maxime, et al.
Published: (2023)
by: Brunswic, Leo Maxime, et al.
Published: (2023)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
by: Wang, Youkang, et al.
Published: (2025)
by: Wang, Youkang, et al.
Published: (2025)
What is the $\textit{intrinsic}$ dimension of your binary data? -- and how to compute it quickly
by: Hanika, Tom, et al.
Published: (2024)
by: Hanika, Tom, et al.
Published: (2024)
Large-scale Urban Facility Location Selection with Knowledge-informed Reinforcement Learning
by: Su, Hongyuan, et al.
Published: (2024)
by: Su, Hongyuan, et al.
Published: (2024)
Refutation of Spectral Graph Theory Conjectures with Search Algorithms)
by: Roucairol, Milo, et al.
Published: (2024)
by: Roucairol, Milo, et al.
Published: (2024)
Unveiling Interesting Insights: Monte Carlo Tree Search for Knowledge Discovery
by: Totis, Pietro, et al.
Published: (2025)
by: Totis, Pietro, et al.
Published: (2025)
In-Context Learning with Topological Information for Knowledge Graph Completion
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
ASNN: Learning to Suggest Neural Architectures from Performance Distributions
by: Hong, Jinwook
Published: (2025)
by: Hong, Jinwook
Published: (2025)
Parameter Tuning of the Firefly Algorithm by Standard Monte Carlo and Quasi-Monte Carlo Methods
by: Joy, Geethu, et al.
Published: (2024)
by: Joy, Geethu, et al.
Published: (2024)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Active Inference with a Self-Prior in the Mirror-Mark Task
by: Kim, Dongmin, et al.
Published: (2026)
by: Kim, Dongmin, et al.
Published: (2026)
DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay
by: Perkins, Daniel, et al.
Published: (2025)
by: Perkins, Daniel, et al.
Published: (2025)
Subset Selection for Fine-Tuning: A Utility-Diversity Balanced Approach for Mathematical Domain Adaptation
by: Kotecha, Madhav, et al.
Published: (2025)
by: Kotecha, Madhav, et al.
Published: (2025)
Maximally Permissive Reward Machines
by: Varricchione, Giovanni, et al.
Published: (2024)
by: Varricchione, Giovanni, et al.
Published: (2024)
A Neural Affinity Framework for Abstract Reasoning: Diagnosing the Compositional Gap in Transformer Architectures via Procedural Task Taxonomy
by: Ingram, Miguel, et al.
Published: (2025)
by: Ingram, Miguel, et al.
Published: (2025)
FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
by: Mohammadzadeh, Saeed, et al.
Published: (2025)
by: Mohammadzadeh, Saeed, et al.
Published: (2025)
Dynamic Observation Policies in Observation Cost-Sensitive Reinforcement Learning
by: Bellinger, Colin, et al.
Published: (2023)
by: Bellinger, Colin, et al.
Published: (2023)
The Boundaries of Verifiable Accuracy, Robustness, and Generalisation in Deep Learning
by: Bastounis, Alexander, et al.
Published: (2023)
by: Bastounis, Alexander, et al.
Published: (2023)
Learning Neural Network Classifiers with Low Model Complexity
by: Jayadeva, et al.
Published: (2017)
by: Jayadeva, et al.
Published: (2017)
Audited Skill-Graph Self-Improvement for Agentic LLMs via Verifiable Rewards, Experience Synthesis, and Continual Memory
by: Huang, Ken, et al.
Published: (2025)
by: Huang, Ken, et al.
Published: (2025)
TNStream: Applying Tightest Neighbors to Micro-Clusters to Define Multi-Density Clusters in Streaming Data
by: Zeng, Qifen, et al.
Published: (2025)
by: Zeng, Qifen, et al.
Published: (2025)
Similar Items
-
From Sparse to Dense: Toddler-inspired Reward Transition in Goal-Oriented Reinforcement Learning
by: Park, Junseok, et al.
Published: (2025) -
Optimizing MoE Routers: Design, Implementation, and Evaluation in Transformer Models
by: Harvey, Daniel Fidel, et al.
Published: (2025) -
From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning
by: Klačan, Ján, et al.
Published: (2026) -
Grouped Sequential Optimization Strategy -- the Application of Hyperparameter Importance Assessment in Deep Learning
by: Wang, Ruinan, et al.
Published: (2025) -
Pushdown Reward Machines for Reinforcement Learning
by: Varricchione, Giovanni, et al.
Published: (2025)