On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
Fuente:
arXiv
Saved in:
| Main Authors: | Shah, Anvay, Anandanarayanan, Ramsundar, Moharir, Sharayu, Kalyanakrishnan, Shivaram |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Observation-Free Attacks on Online Learning to Rank
by: Chattopadhyay, Sameep, et al.
Published: (2025)
by: Chattopadhyay, Sameep, et al.
Published: (2025)
Influencing Bandits: Arm Selection for Preference Shaping
by: Nadkarni, Viraj, et al.
Published: (2024)
by: Nadkarni, Viraj, et al.
Published: (2024)
Fixed-Budget Constrained Best Arm Identification in Grouped Bandits
by: Mukherjee, Raunak, et al.
Published: (2026)
by: Mukherjee, Raunak, et al.
Published: (2026)
Efficient Computation of Blackwell Optimal Policies using Rational Functions
by: Mukherjee, Dibyangshu, et al.
Published: (2025)
by: Mukherjee, Dibyangshu, et al.
Published: (2025)
Howard's Policy Iteration is Subexponential for Deterministic Markov Decision Problems with Rewards of Fixed Bit-size and Arbitrary Discount Factor
by: Mukherjee, Dibyangshu, et al.
Published: (2025)
by: Mukherjee, Dibyangshu, et al.
Published: (2025)
A View of the Certainty-Equivalence Method for PAC RL as an Application of the Trajectory Tree Method
by: Kalyanakrishnan, Shivaram, et al.
Published: (2025)
by: Kalyanakrishnan, Shivaram, et al.
Published: (2025)
Cascading Bandits With Feedback
by: Prakash, R Sri, et al.
Published: (2025)
by: Prakash, R Sri, et al.
Published: (2025)
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
by: Liu, Xutong, et al.
Published: (2023)
by: Liu, Xutong, et al.
Published: (2023)
Unreliable Multi-Armed Bandits: A Novel Approach to Recommendation Systems
by: Ravi, Aditya Narayan, et al.
Published: (2019)
by: Ravi, Aditya Narayan, et al.
Published: (2019)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
by: Banerjee, Debangshu, et al.
Published: (2023)
by: Banerjee, Debangshu, et al.
Published: (2023)
Batch-Size Independent Regret Bounds for Combinatorial Semi-Bandits with Probabilistically Triggered Arms or Independent Arms
by: Liu, Xutong, et al.
Published: (2022)
by: Liu, Xutong, et al.
Published: (2022)
Constrained Best Arm Identification in Grouped Bandits
by: Dharod, Sahil, et al.
Published: (2024)
by: Dharod, Sahil, et al.
Published: (2024)
Fairness for Workers Who Pull the Arms: An Index Based Policy for Allocation of Restless Bandit Tasks
by: Biswas, Arpita, et al.
Published: (2023)
by: Biswas, Arpita, et al.
Published: (2023)
Learning Policy Committees for Effective Personalization in MDPs with Diverse Tasks
by: Ge, Luise, et al.
Published: (2025)
by: Ge, Luise, et al.
Published: (2025)
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023)
by: Agnihotri, Akhil, et al.
Published: (2023)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Lagrangian Index Policy for Restless Bandits with Average Reward
by: Avrachenkov, Konstantin, et al.
Published: (2024)
by: Avrachenkov, Konstantin, et al.
Published: (2024)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024)
by: Bai, Qinbo, et al.
Published: (2024)
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
by: Shimizu, Tatsuhiro, et al.
Published: (2024)
by: Shimizu, Tatsuhiro, et al.
Published: (2024)
Efficient Solution and Learning of Robust Factored MDPs
by: Schnitzer, Yannik, et al.
Published: (2025)
by: Schnitzer, Yannik, et al.
Published: (2025)
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
by: Manivannan, Sanjeev, et al.
Published: (2026)
by: Manivannan, Sanjeev, et al.
Published: (2026)
Optimizing Vehicular Networks with Variational Quantum Circuits-based Reinforcement Learning
by: Yan, Zijiang, et al.
Published: (2024)
by: Yan, Zijiang, et al.
Published: (2024)
Identifying All ε-Best Arms in (Misspecified) Linear Bandits
by: Li, Zhekai, et al.
Published: (2025)
by: Li, Zhekai, et al.
Published: (2025)
Teaching Precommitted Agents: Model-Free Policy Evaluation and Control in Quasi-Hyperbolic Discounted MDPs
by: Eshwar, S. R.
Published: (2025)
by: Eshwar, S. R.
Published: (2025)
Low-Regret and Low-Complexity Learning for Hierarchical Inference
by: Chattopadhyay, Sameep, et al.
Published: (2025)
by: Chattopadhyay, Sameep, et al.
Published: (2025)
Optimizing Warfarin Dosing Using Contextual Bandit: An Offline Policy Learning and Evaluation Method
by: Huang, Yong, et al.
Published: (2024)
by: Huang, Yong, et al.
Published: (2024)
Model-Free Learning and Optimal Policy Design in Multi-Agent MDPs Under Probabilistic Agent Dropout
by: Fiscko, Carmel, et al.
Published: (2023)
by: Fiscko, Carmel, et al.
Published: (2023)
Retrosynthesis Planning via Worst-path Policy Optimisation in Tree-structured MDPs
by: Wang, Mianchu, et al.
Published: (2025)
by: Wang, Mianchu, et al.
Published: (2025)
DeepAveragers: Offline Reinforcement Learning by Solving Derived Non-Parametric MDPs
by: Shrestha, Aayam, et al.
Published: (2020)
by: Shrestha, Aayam, et al.
Published: (2020)
Projection by Convolution: Optimal Sample Complexity for Reinforcement Learning in Continuous-Space MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
by: Mhammedi, Zakaria
Published: (2024)
by: Mhammedi, Zakaria
Published: (2024)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
Tree Ensembles for Contextual Bandits
by: Nilsson, Hannes, et al.
Published: (2024)
by: Nilsson, Hannes, et al.
Published: (2024)
Using Common Random Numbers for Simulation-based Planning with Rollouts
by: Yadav, Sandarbh, et al.
Published: (2026)
by: Yadav, Sandarbh, et al.
Published: (2026)
Open-Source Molecular Processing Pipeline for Generating Molecules
by: Shreyas, V, et al.
Published: (2024)
by: Shreyas, V, et al.
Published: (2024)
Low-Rank MDPs with Continuous Action Spaces
by: Bennett, Andrew, et al.
Published: (2023)
by: Bennett, Andrew, et al.
Published: (2023)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
Adaptive Opponent Policy Detection in Multi-Agent MDPs: Real-Time Strategy Switch Identification Using Running Error Estimation
by: Mridul, Mohidul Haque, et al.
Published: (2024)
by: Mridul, Mohidul Haque, et al.
Published: (2024)
Similar Items
-
Observation-Free Attacks on Online Learning to Rank
by: Chattopadhyay, Sameep, et al.
Published: (2025) -
Influencing Bandits: Arm Selection for Preference Shaping
by: Nadkarni, Viraj, et al.
Published: (2024) -
Fixed-Budget Constrained Best Arm Identification in Grouped Bandits
by: Mukherjee, Raunak, et al.
Published: (2026) -
Efficient Computation of Blackwell Optimal Policies using Rational Functions
by: Mukherjee, Dibyangshu, et al.
Published: (2025) -
Howard's Policy Iteration is Subexponential for Deterministic Markov Decision Problems with Rewards of Fixed Bit-size and Arbitrary Discount Factor
by: Mukherjee, Dibyangshu, et al.
Published: (2025)