Gespeichert in:
| Hauptverfasser: | Nguyen, Tuan Ngo, Barrett, Jay, Jun, Kwang-Sung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2411.00405 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
$\varepsilon$-Good Action Identification in Fixed-Budget Monte Carlo Tree Search
von: Li, Yinan, et al.
Veröffentlicht: (2026)
von: Li, Yinan, et al.
Veröffentlicht: (2026)
Fixing the Loose Brake: Exponential-Tailed Stopping Time in Best Arm Identification
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits
von: Jang, Kyoungseok, et al.
Veröffentlicht: (2024)
von: Jang, Kyoungseok, et al.
Veröffentlicht: (2024)
Nearly Optimal Active Preference Learning and Its Application to LLM Alignment
von: Zhao, Yao, et al.
Veröffentlicht: (2026)
von: Zhao, Yao, et al.
Veröffentlicht: (2026)
Instance-Dependent Continuous-Time Reinforcement Learning via Maximum Likelihood Estimation
von: Zhao, Runze, et al.
Veröffentlicht: (2025)
von: Zhao, Runze, et al.
Veröffentlicht: (2025)
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss
von: Li, Yinan, et al.
Veröffentlicht: (2025)
von: Li, Yinan, et al.
Veröffentlicht: (2025)
Improve Value Estimation of Q Function and Reshape Reward with Monte Carlo Tree Search
von: Li, Jiamian
Veröffentlicht: (2024)
von: Li, Jiamian
Veröffentlicht: (2024)
Noise-Adaptive Confidence Sets for Linear Bandits and Application to Bayesian Optimization
von: Jun, Kwang-Sung, et al.
Veröffentlicht: (2024)
von: Jun, Kwang-Sung, et al.
Veröffentlicht: (2024)
An Upper Confidence Bound Approach to Estimating the Maximum Mean
von: Kun, Zhang, et al.
Veröffentlicht: (2024)
von: Kun, Zhang, et al.
Veröffentlicht: (2024)
GL-LowPopArt: A Nearly Instance-Wise Minimax-Optimal Estimator for Generalized Low-Rank Trace Regression
von: Lee, Junghyun, et al.
Veröffentlicht: (2025)
von: Lee, Junghyun, et al.
Veröffentlicht: (2025)
Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards
von: Qin, Hao, et al.
Veröffentlicht: (2023)
von: Qin, Hao, et al.
Veröffentlicht: (2023)
Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
UNSAT Solver Synthesis via Monte Carlo Forest Search
von: Cameron, Chris, et al.
Veröffentlicht: (2022)
von: Cameron, Chris, et al.
Veröffentlicht: (2022)
Better-than-KL PAC-Bayes Bounds
von: Kuzborskij, Ilja, et al.
Veröffentlicht: (2024)
von: Kuzborskij, Ilja, et al.
Veröffentlicht: (2024)
Epistemic Monte Carlo Tree Search
von: Oren, Yaniv, et al.
Veröffentlicht: (2022)
von: Oren, Yaniv, et al.
Veröffentlicht: (2022)
A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits
von: Lee, Junghyun, et al.
Veröffentlicht: (2024)
von: Lee, Junghyun, et al.
Veröffentlicht: (2024)
Enhancing Bayesian Network Structural Learning with Monte Carlo Tree Search
von: Laborda, Jorge D., et al.
Veröffentlicht: (2025)
von: Laborda, Jorge D., et al.
Veröffentlicht: (2025)
MEC: Machine-Learning-Assisted Generalized Entropy Calibration for Semi-Supervised Mean Estimation
von: Lee, Se Yoon, et al.
Veröffentlicht: (2026)
von: Lee, Se Yoon, et al.
Veröffentlicht: (2026)
Dependency-aware Maximum Likelihood Estimation for Active Learning
von: Kalkanli, Beyza, et al.
Veröffentlicht: (2025)
von: Kalkanli, Beyza, et al.
Veröffentlicht: (2025)
Fixed Budget is No Harder Than Fixed Confidence in Best-Arm Identification up to Logarithmic Factors
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2026)
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2026)
Entropic Risk-Aware Monte Carlo Tree Search
von: Santos, Pedro P., et al.
Veröffentlicht: (2026)
von: Santos, Pedro P., et al.
Veröffentlicht: (2026)
An Efficient Algorithm for Thresholding Monte Carlo Tree Search
von: Nameki, Shoma, et al.
Veröffentlicht: (2026)
von: Nameki, Shoma, et al.
Veröffentlicht: (2026)
Improving Monte Carlo Tree Search for Symbolic Regression
von: Huang, Zhengyao, et al.
Veröffentlicht: (2025)
von: Huang, Zhengyao, et al.
Veröffentlicht: (2025)
Minimum Empirical Divergence for Sub-Gaussian Linear Bandits
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024)
Power Mean Estimation in Stochastic Monte-Carlo Tree_Search
von: Dam, Tuan, et al.
Veröffentlicht: (2024)
von: Dam, Tuan, et al.
Veröffentlicht: (2024)
Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing
von: Ryu, J. Jon, et al.
Veröffentlicht: (2025)
von: Ryu, J. Jon, et al.
Veröffentlicht: (2025)
Twice Sequential Monte Carlo for Tree Search
von: Oren, Yaniv, et al.
Veröffentlicht: (2025)
von: Oren, Yaniv, et al.
Veröffentlicht: (2025)
Monte Carlo Tree Search with Boltzmann Exploration
von: Painter, Michael, et al.
Veröffentlicht: (2024)
von: Painter, Michael, et al.
Veröffentlicht: (2024)
Doubly Robust Monte Carlo Tree Search
von: Liu, Manqing, et al.
Veröffentlicht: (2025)
von: Liu, Manqing, et al.
Veröffentlicht: (2025)
Private Gradient Descent for Linear Regression: Tighter Error Bounds and Instance-Specific Uncertainty Estimation
von: Brown, Gavin, et al.
Veröffentlicht: (2024)
von: Brown, Gavin, et al.
Veröffentlicht: (2024)
Monte Carlo Permutation Search
von: Cazenave, Tristan
Veröffentlicht: (2025)
von: Cazenave, Tristan
Veröffentlicht: (2025)
Instance-Dependent Regret Bounds for Nonstochastic Linear Partial Monitoring
von: Di Gennaro, Federico, et al.
Veröffentlicht: (2025)
von: Di Gennaro, Federico, et al.
Veröffentlicht: (2025)
Active Level Set Estimation for Continuous Search Space with Theoretical Guarantee
von: Ngo, Giang, et al.
Veröffentlicht: (2024)
von: Ngo, Giang, et al.
Veröffentlicht: (2024)
Gap-Dependent Bounds for Federated $Q$-learning
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
von: Zhang, Haochen, et al.
Veröffentlicht: (2025)
Can Transformers Learn to Verify During Backtracking Search?
von: Phua, Yin Jun, et al.
Veröffentlicht: (2026)
von: Phua, Yin Jun, et al.
Veröffentlicht: (2026)
Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition
von: Zheng, Zhong, et al.
Veröffentlicht: (2024)
von: Zheng, Zhong, et al.
Veröffentlicht: (2024)
Anytime Sequential Halving in Monte-Carlo Tree Search
von: Sagers, Dominic, et al.
Veröffentlicht: (2024)
von: Sagers, Dominic, et al.
Veröffentlicht: (2024)
Monte Carlo Tree Search in the Presence of Transition Uncertainty
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2023)
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2023)
Improving GFlowNets with Monte Carlo Tree Search
von: Morozov, Nikita, et al.
Veröffentlicht: (2024)
von: Morozov, Nikita, et al.
Veröffentlicht: (2024)
Robust Transfer Learning for Active Level Set Estimation with Locally Adaptive Gaussian Process Prior
von: Ngo, Giang, et al.
Veröffentlicht: (2024)
von: Ngo, Giang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
$\varepsilon$-Good Action Identification in Fixed-Budget Monte Carlo Tree Search
von: Li, Yinan, et al.
Veröffentlicht: (2026) -
Fixing the Loose Brake: Exponential-Tailed Stopping Time in Best Arm Identification
von: Balagopalan, Kapilan, et al.
Veröffentlicht: (2024) -
Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits
von: Jang, Kyoungseok, et al.
Veröffentlicht: (2024) -
Nearly Optimal Active Preference Learning and Its Application to LLM Alignment
von: Zhao, Yao, et al.
Veröffentlicht: (2026) -
Instance-Dependent Continuous-Time Reinforcement Learning via Maximum Likelihood Estimation
von: Zhao, Runze, et al.
Veröffentlicht: (2025)