HAVER: Instance-Dependent Error Bounds for Maximum Mean Estimation and Applications to Q-Learning and Monte Carlo Tree Search
Fuente:
arXiv
Salvato in:
| Autori principali: | Nguyen, Tuan Ngo, Barrett, Jay, Jun, Kwang-Sung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
$\varepsilon$-Good Action Identification in Fixed-Budget Monte Carlo Tree Search
di: Li, Yinan, et al.
Pubblicazione: (2026)
di: Li, Yinan, et al.
Pubblicazione: (2026)
Fixing the Loose Brake: Exponential-Tailed Stopping Time in Best Arm Identification
di: Balagopalan, Kapilan, et al.
Pubblicazione: (2024)
di: Balagopalan, Kapilan, et al.
Pubblicazione: (2024)
Instance-Dependent Continuous-Time Reinforcement Learning via Maximum Likelihood Estimation
di: Zhao, Runze, et al.
Pubblicazione: (2025)
di: Zhao, Runze, et al.
Pubblicazione: (2025)
Nearly Optimal Active Preference Learning and Its Application to LLM Alignment
di: Zhao, Yao, et al.
Pubblicazione: (2026)
di: Zhao, Yao, et al.
Pubblicazione: (2026)
Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits
di: Jang, Kyoungseok, et al.
Pubblicazione: (2024)
di: Jang, Kyoungseok, et al.
Pubblicazione: (2024)
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss
di: Li, Yinan, et al.
Pubblicazione: (2025)
di: Li, Yinan, et al.
Pubblicazione: (2025)
Improve Value Estimation of Q Function and Reshape Reward with Monte Carlo Tree Search
di: Li, Jiamian
Pubblicazione: (2024)
di: Li, Jiamian
Pubblicazione: (2024)
An Upper Confidence Bound Approach to Estimating the Maximum Mean
di: Kun, Zhang, et al.
Pubblicazione: (2024)
di: Kun, Zhang, et al.
Pubblicazione: (2024)
Noise-Adaptive Confidence Sets for Linear Bandits and Application to Bayesian Optimization
di: Jun, Kwang-Sung, et al.
Pubblicazione: (2024)
di: Jun, Kwang-Sung, et al.
Pubblicazione: (2024)
GL-LowPopArt: A Nearly Instance-Wise Minimax-Optimal Estimator for Generalized Low-Rank Trace Regression
di: Lee, Junghyun, et al.
Pubblicazione: (2025)
di: Lee, Junghyun, et al.
Pubblicazione: (2025)
Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards
di: Qin, Hao, et al.
Pubblicazione: (2023)
di: Qin, Hao, et al.
Pubblicazione: (2023)
Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion
di: Lee, Junghyun, et al.
Pubblicazione: (2023)
di: Lee, Junghyun, et al.
Pubblicazione: (2023)
Enhancing Bayesian Network Structural Learning with Monte Carlo Tree Search
di: Laborda, Jorge D., et al.
Pubblicazione: (2025)
di: Laborda, Jorge D., et al.
Pubblicazione: (2025)
Dependency-aware Maximum Likelihood Estimation for Active Learning
di: Kalkanli, Beyza, et al.
Pubblicazione: (2025)
di: Kalkanli, Beyza, et al.
Pubblicazione: (2025)
Epistemic Monte Carlo Tree Search
di: Oren, Yaniv, et al.
Pubblicazione: (2022)
di: Oren, Yaniv, et al.
Pubblicazione: (2022)
UNSAT Solver Synthesis via Monte Carlo Forest Search
di: Cameron, Chris, et al.
Pubblicazione: (2022)
di: Cameron, Chris, et al.
Pubblicazione: (2022)
Better-than-KL PAC-Bayes Bounds
di: Kuzborskij, Ilja, et al.
Pubblicazione: (2024)
di: Kuzborskij, Ilja, et al.
Pubblicazione: (2024)
MEC: Machine-Learning-Assisted Generalized Entropy Calibration for Semi-Supervised Mean Estimation
di: Lee, Se Yoon, et al.
Pubblicazione: (2026)
di: Lee, Se Yoon, et al.
Pubblicazione: (2026)
Entropic Risk-Aware Monte Carlo Tree Search
di: Santos, Pedro P., et al.
Pubblicazione: (2026)
di: Santos, Pedro P., et al.
Pubblicazione: (2026)
An Efficient Algorithm for Thresholding Monte Carlo Tree Search
di: Nameki, Shoma, et al.
Pubblicazione: (2026)
di: Nameki, Shoma, et al.
Pubblicazione: (2026)
Improving Monte Carlo Tree Search for Symbolic Regression
di: Huang, Zhengyao, et al.
Pubblicazione: (2025)
di: Huang, Zhengyao, et al.
Pubblicazione: (2025)
A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits
di: Lee, Junghyun, et al.
Pubblicazione: (2024)
di: Lee, Junghyun, et al.
Pubblicazione: (2024)
Private Gradient Descent for Linear Regression: Tighter Error Bounds and Instance-Specific Uncertainty Estimation
di: Brown, Gavin, et al.
Pubblicazione: (2024)
di: Brown, Gavin, et al.
Pubblicazione: (2024)
Instance-Dependent Regret Bounds for Nonstochastic Linear Partial Monitoring
di: Di Gennaro, Federico, et al.
Pubblicazione: (2025)
di: Di Gennaro, Federico, et al.
Pubblicazione: (2025)
Twice Sequential Monte Carlo for Tree Search
di: Oren, Yaniv, et al.
Pubblicazione: (2025)
di: Oren, Yaniv, et al.
Pubblicazione: (2025)
Monte Carlo Tree Search with Boltzmann Exploration
di: Painter, Michael, et al.
Pubblicazione: (2024)
di: Painter, Michael, et al.
Pubblicazione: (2024)
Doubly Robust Monte Carlo Tree Search
di: Liu, Manqing, et al.
Pubblicazione: (2025)
di: Liu, Manqing, et al.
Pubblicazione: (2025)
Gap-Dependent Bounds for Federated $Q$-learning
di: Zhang, Haochen, et al.
Pubblicazione: (2025)
di: Zhang, Haochen, et al.
Pubblicazione: (2025)
Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition
di: Zheng, Zhong, et al.
Pubblicazione: (2024)
di: Zheng, Zhong, et al.
Pubblicazione: (2024)
Minimum Empirical Divergence for Sub-Gaussian Linear Bandits
di: Balagopalan, Kapilan, et al.
Pubblicazione: (2024)
di: Balagopalan, Kapilan, et al.
Pubblicazione: (2024)
Power Mean Estimation in Stochastic Monte-Carlo Tree_Search
di: Dam, Tuan, et al.
Pubblicazione: (2024)
di: Dam, Tuan, et al.
Pubblicazione: (2024)
Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing
di: Ryu, J. Jon, et al.
Pubblicazione: (2025)
di: Ryu, J. Jon, et al.
Pubblicazione: (2025)
Anytime Sequential Halving in Monte-Carlo Tree Search
di: Sagers, Dominic, et al.
Pubblicazione: (2024)
di: Sagers, Dominic, et al.
Pubblicazione: (2024)
Monte Carlo Tree Search in the Presence of Transition Uncertainty
di: Kohankhaki, Farnaz, et al.
Pubblicazione: (2023)
di: Kohankhaki, Farnaz, et al.
Pubblicazione: (2023)
Improving GFlowNets with Monte Carlo Tree Search
di: Morozov, Nikita, et al.
Pubblicazione: (2024)
di: Morozov, Nikita, et al.
Pubblicazione: (2024)
Monte Carlo Permutation Search
di: Cazenave, Tristan
Pubblicazione: (2025)
di: Cazenave, Tristan
Pubblicazione: (2025)
Active Level Set Estimation for Continuous Search Space with Theoretical Guarantee
di: Ngo, Giang, et al.
Pubblicazione: (2024)
di: Ngo, Giang, et al.
Pubblicazione: (2024)
Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
di: Xie, Yuxi, et al.
Pubblicazione: (2024)
di: Xie, Yuxi, et al.
Pubblicazione: (2024)
Robust Transfer Learning for Active Level Set Estimation with Locally Adaptive Gaussian Process Prior
di: Ngo, Giang, et al.
Pubblicazione: (2024)
di: Ngo, Giang, et al.
Pubblicazione: (2024)
Fixed Budget is No Harder Than Fixed Confidence in Best-Arm Identification up to Logarithmic Factors
di: Balagopalan, Kapilan, et al.
Pubblicazione: (2026)
di: Balagopalan, Kapilan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
$\varepsilon$-Good Action Identification in Fixed-Budget Monte Carlo Tree Search
di: Li, Yinan, et al.
Pubblicazione: (2026) -
Fixing the Loose Brake: Exponential-Tailed Stopping Time in Best Arm Identification
di: Balagopalan, Kapilan, et al.
Pubblicazione: (2024) -
Instance-Dependent Continuous-Time Reinforcement Learning via Maximum Likelihood Estimation
di: Zhao, Runze, et al.
Pubblicazione: (2025) -
Nearly Optimal Active Preference Learning and Its Application to LLM Alignment
di: Zhao, Yao, et al.
Pubblicazione: (2026) -
Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits
di: Jang, Kyoungseok, et al.
Pubblicazione: (2024)