Super-Exponential Regret for UCT, AlphaGo and Variants
Fuente:
arXiv
Saved in:
| Main Authors: | Orseau, Laurent, Munos, Remi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AlphaGo Moment for Model Architecture Discovery
by: Liu, Yixiu, et al.
Published: (2025)
by: Liu, Yixiu, et al.
Published: (2025)
Isotuning With Applications To Scale-Free Online Learning
by: Orseau, Laurent, et al.
Published: (2021)
by: Orseau, Laurent, et al.
Published: (2021)
Levin Tree Search with Context Models
by: Orseau, Laurent, et al.
Published: (2023)
by: Orseau, Laurent, et al.
Published: (2023)
Regret-Guided Search Control for Efficient Learning in AlphaZero
by: Tsai, Yun-Jui, et al.
Published: (2026)
by: Tsai, Yun-Jui, et al.
Published: (2026)
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Exponential Speedups by Rerooting Levin Tree Search
by: Orseau, Laurent, et al.
Published: (2024)
by: Orseau, Laurent, et al.
Published: (2024)
Temporal Difference Flows
by: Farebrother, Jesse, et al.
Published: (2025)
by: Farebrother, Jesse, et al.
Published: (2025)
Understanding Prompt Tuning and In-Context Learning via Meta-Learning
by: Genewein, Tim, et al.
Published: (2025)
by: Genewein, Tim, et al.
Published: (2025)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
Affordances Enable Partial World Modeling with LLMs
by: Khetarpal, Khimya, et al.
Published: (2026)
by: Khetarpal, Khimya, et al.
Published: (2026)
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
by: Wu, Ti-Rong, et al.
Published: (2023)
by: Wu, Ti-Rong, et al.
Published: (2023)
Finding Increasingly Large Extremal Graphs with AlphaZero and Tabu Search
by: Mehrabian, Abbas, et al.
Published: (2023)
by: Mehrabian, Abbas, et al.
Published: (2023)
Super Level Sets and Exponential Decay: A Synergistic Approach to Stable Neural Network Training
by: Chaudhary, Jatin, et al.
Published: (2024)
by: Chaudhary, Jatin, et al.
Published: (2024)
Reasoning without Regret
by: Chitra, Tarun
Published: (2025)
by: Chitra, Tarun
Published: (2025)
AlphaZero-Edu: Democratizing Access to AlphaZero
by: Li, Ruitong, et al.
Published: (2025)
by: Li, Ruitong, et al.
Published: (2025)
No Regrets: Investigating and Improving Regret Approximations for Curriculum Discovery
by: Rutherford, Alexander, et al.
Published: (2024)
by: Rutherford, Alexander, et al.
Published: (2024)
Learning Universal Predictors
by: Grau-Moya, Jordi, et al.
Published: (2024)
by: Grau-Moya, Jordi, et al.
Published: (2024)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Generalized Preference Optimization: A Unified Approach to Offline Alignment
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Understanding the performance gap between online and offline alignment algorithms
by: Tang, Yunhao, et al.
Published: (2024)
by: Tang, Yunhao, et al.
Published: (2024)
Regret-Free Reinforcement Learning for LTL Specifications
by: Majumdar, Rupak, et al.
Published: (2024)
by: Majumdar, Rupak, et al.
Published: (2024)
Refining Minimax Regret for Unsupervised Environment Design
by: Beukman, Michael, et al.
Published: (2024)
by: Beukman, Michael, et al.
Published: (2024)
Regret-Based Defense in Adversarial Reinforcement Learning
by: Belaire, Roman, et al.
Published: (2023)
by: Belaire, Roman, et al.
Published: (2023)
A Regret Perspective on Online Multiple Testing
by: Hao, Qingyang, et al.
Published: (2026)
by: Hao, Qingyang, et al.
Published: (2026)
Optimizing Return Distributions with Distributional Dynamic Programming
by: Pires, Bernardo Ávila, et al.
Published: (2025)
by: Pires, Bernardo Ávila, et al.
Published: (2025)
Provably Efficient Exploration in Reward Machines with Low Regret
by: Bourel, Hippolyte, et al.
Published: (2024)
by: Bourel, Hippolyte, et al.
Published: (2024)
Efficient Skill Discovery via Regret-Aware Optimization
by: Zhang, He, et al.
Published: (2025)
by: Zhang, He, et al.
Published: (2025)
Data-Driven Online Model Selection With Regret Guarantees
by: Pacchiano, Aldo, et al.
Published: (2023)
by: Pacchiano, Aldo, et al.
Published: (2023)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
by: He, Jiafan, et al.
Published: (2025)
by: He, Jiafan, et al.
Published: (2025)
Regret-Based Federated Causal Discovery with Unknown Interventions
by: Baldo, Federico, et al.
Published: (2025)
by: Baldo, Federico, et al.
Published: (2025)
Kernelized Reinforcement Learning with Order Optimal Regret Bounds
by: Vakili, Sattar, et al.
Published: (2023)
by: Vakili, Sattar, et al.
Published: (2023)
APO: Alpha-Divergence Preference Optimization
by: Zixian, Wang
Published: (2025)
by: Zixian, Wang
Published: (2025)
Joint Embeddings Go Temporal
by: Ennadir, Sofiane, et al.
Published: (2025)
by: Ennadir, Sofiane, et al.
Published: (2025)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Adversarial Environment Design via Regret-Guided Diffusion Models
by: Chung, Hojun, et al.
Published: (2024)
by: Chung, Hojun, et al.
Published: (2024)
PROWL: Prioritized Regret-Driven Optimization for World Model Learning
by: Güzel, Ahmet H., et al.
Published: (2026)
by: Güzel, Ahmet H., et al.
Published: (2026)
Improved Bayesian Regret Bounds for Thompson Sampling in Reinforcement Learning
by: Moradipari, Ahmadreza, et al.
Published: (2023)
by: Moradipari, Ahmadreza, et al.
Published: (2023)
$(ε, u)$-Adaptive Regret Minimization in Heavy-Tailed Bandits
by: Genalti, Gianmarco, et al.
Published: (2023)
by: Genalti, Gianmarco, et al.
Published: (2023)
Regret Bounds and Reinforcement Learning Exploration of EXP-based Algorithms
by: Xu, Mengfan, et al.
Published: (2020)
by: Xu, Mengfan, et al.
Published: (2020)
Toward Optimal Regret in Robust Pricing: Decoupling Corruption and Time
by: Kalupahana, Kalana, et al.
Published: (2026)
by: Kalupahana, Kalana, et al.
Published: (2026)
Similar Items
-
AlphaGo Moment for Model Architecture Discovery
by: Liu, Yixiu, et al.
Published: (2025) -
Isotuning With Applications To Scale-Free Online Learning
by: Orseau, Laurent, et al.
Published: (2021) -
Levin Tree Search with Context Models
by: Orseau, Laurent, et al.
Published: (2023) -
Regret-Guided Search Control for Efficient Learning in AlphaZero
by: Tsai, Yun-Jui, et al.
Published: (2026) -
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)