Representative Action Selection for Large Action Space: From Bandits to MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Quan, Mannor, Shie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Representative Action Selection for Large Action Space Bandit Families
by: Zhou, Quan, et al.
Published: (2025)
by: Zhou, Quan, et al.
Published: (2025)
Bandit Allocational Instability
by: Chen, Yilun, et al.
Published: (2026)
by: Chen, Yilun, et al.
Published: (2026)
Model Predictive Control is Almost Optimal for Restless Bandit
by: Gast, Nicolas, et al.
Published: (2024)
by: Gast, Nicolas, et al.
Published: (2024)
Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits
by: Narasimha, Dheeraj, et al.
Published: (2025)
by: Narasimha, Dheeraj, et al.
Published: (2025)
Landscape of Policy Optimization for Finite Horizon MDPs with General State and Action
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
Early Stopping in Contextual Bandits and Inferences
by: Cui, Zihan
Published: (2025)
by: Cui, Zihan
Published: (2025)
Lagrangian Index Policy for Restless Bandits with Average Reward
by: Avrachenkov, Konstantin, et al.
Published: (2024)
by: Avrachenkov, Konstantin, et al.
Published: (2024)
Convergence of Actor-Critic Learning for Mean Field Games and Mean Field Control in Continuous Spaces
by: Fouque, Jean-Pierre, et al.
Published: (2025)
by: Fouque, Jean-Pierre, et al.
Published: (2025)
Regret of exploratory policy improvement and $q$-learning
by: Tang, Wenpin, et al.
Published: (2024)
by: Tang, Wenpin, et al.
Published: (2024)
Large Deviation Upper Bounds and Improved MSE Rates of Nonlinear SGD: Heavy-tailed Noise and Power of Symmetry
by: Armacki, Aleksandar, et al.
Published: (2024)
by: Armacki, Aleksandar, et al.
Published: (2024)
Neural Hilbert Ladders: Multi-Layer Neural Networks in Function Space
by: Chen, Zhengdao
Published: (2023)
by: Chen, Zhengdao
Published: (2023)
Multi-Action Restless Bandits with Weakly Coupled Constraints: Simultaneous Learning and Control
by: Fu, Jing, et al.
Published: (2024)
by: Fu, Jing, et al.
Published: (2024)
Non-convex entropic mean-field optimization via Best Response flow
by: Lascu, Razvan-Andrei, et al.
Published: (2025)
by: Lascu, Razvan-Andrei, et al.
Published: (2025)
ODE approximation for the Adam algorithm: General and overparametrized setting
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Flatness-Aware Stochastic Gradient Langevin Dynamics
by: Bruno, Stefano, et al.
Published: (2025)
by: Bruno, Stefano, et al.
Published: (2025)
Admission Control of Quasi-Reversible Queueing Systems: Optimization and Reinforcement Learning
by: Comte, Céline, et al.
Published: (2025)
by: Comte, Céline, et al.
Published: (2025)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
by: Kassing, Sebastian, et al.
Published: (2025)
by: Kassing, Sebastian, et al.
Published: (2025)
Designing Algorithms for Entropic Optimal Transport from an Optimisation Perspective
by: Srinivasan, Vishwak, et al.
Published: (2025)
by: Srinivasan, Vishwak, et al.
Published: (2025)
Benchmarking Diffusion Annealing-Based Bayesian Inverse Problem Solvers
by: Crafts, Evan Scope, et al.
Published: (2025)
by: Crafts, Evan Scope, et al.
Published: (2025)
A stochastic gradient descent algorithm with random search directions
by: Gbaguidi, Eméric
Published: (2025)
by: Gbaguidi, Eméric
Published: (2025)
Algorithmic Stability of Stochastic Gradient Descent with Momentum under Heavy-Tailed Noise
by: Dang, Thanh, et al.
Published: (2025)
by: Dang, Thanh, et al.
Published: (2025)
Reinforcement Learning with Random Time Horizons
by: Borrell, Enric Ribera, et al.
Published: (2025)
by: Borrell, Enric Ribera, et al.
Published: (2025)
A Distributional View of High Dimensional Optimization
by: Benning, Felix
Published: (2025)
by: Benning, Felix
Published: (2025)
Improved Approximation Algorithms for Orthogonally Constrained Problems Using Semidefinite Optimization
by: Cory-Wright, Ryan, et al.
Published: (2025)
by: Cory-Wright, Ryan, et al.
Published: (2025)
When Machine Learning Meets Importance Sampling: A More Efficient Rare Event Estimation Approach
by: Zhao, Ruoning, et al.
Published: (2025)
by: Zhao, Ruoning, et al.
Published: (2025)
Wasserstein Convergence of Score-based Generative Models under Semiconvexity and Discontinuous Gradients
by: Bruno, Stefano, et al.
Published: (2025)
by: Bruno, Stefano, et al.
Published: (2025)
A Generalization Result for Convergence in Learning-to-Optimize
by: Sucker, Michael, et al.
Published: (2024)
by: Sucker, Michael, et al.
Published: (2024)
Properties of Discrete Sliced Wasserstein Losses
by: Tanguy, Eloi, et al.
Published: (2023)
by: Tanguy, Eloi, et al.
Published: (2023)
Decoupled Functional Central Limit Theorems for Two-Time-Scale Stochastic Approximation
by: Han, Yuze, et al.
Published: (2024)
by: Han, Yuze, et al.
Published: (2024)
On propagation of chaos for the Fisher-Rao gradient flow in entropic mean-field optimization
by: Lazić, Petra, et al.
Published: (2026)
by: Lazić, Petra, et al.
Published: (2026)
Value Mirror Descent for Reinforcement Learning
by: Jia, Zhichao, et al.
Published: (2026)
by: Jia, Zhichao, et al.
Published: (2026)
Improved sampling via learned diffusions
by: Richter, Lorenz, et al.
Published: (2023)
by: Richter, Lorenz, et al.
Published: (2023)
Function approximation by neural nets in the mean-field regime: Entropic regularization and controlled McKean-Vlasov dynamics
by: Tzen, Belinda, et al.
Published: (2020)
by: Tzen, Belinda, et al.
Published: (2020)
Asymptotic regularity of a generalised stochastic Halpern scheme
by: Pischke, Nicholas, et al.
Published: (2024)
by: Pischke, Nicholas, et al.
Published: (2024)
Generalized Wasserstein Flow Matching: Transport Plans, Everywhere, All at Once
by: Piening, Moritz, et al.
Published: (2026)
by: Piening, Moritz, et al.
Published: (2026)
Concentration of General Stochastic Approximation Under Heavy-Tailed Markovian Noise
by: Agrawal, Shubhada, et al.
Published: (2026)
by: Agrawal, Shubhada, et al.
Published: (2026)
Approximation and interpolation of deep neural networks
by: Constantinescu, Vlad-Raul, et al.
Published: (2023)
by: Constantinescu, Vlad-Raul, et al.
Published: (2023)
Hilbert's projective metric for functions of bounded growth and exponential convergence of Sinkhorn's algorithm
by: Eckstein, Stephan
Published: (2023)
by: Eckstein, Stephan
Published: (2023)
Posterior Sampling Based on Gradient Flows of the MMD with Negative Distance Kernel
by: Hagemann, Paul, et al.
Published: (2023)
by: Hagemann, Paul, et al.
Published: (2023)
Accelerating Distributed Stochastic Optimization via Self-Repellent Random Walks
by: Hu, Jie, et al.
Published: (2024)
by: Hu, Jie, et al.
Published: (2024)
Similar Items
-
Representative Action Selection for Large Action Space Bandit Families
by: Zhou, Quan, et al.
Published: (2025) -
Bandit Allocational Instability
by: Chen, Yilun, et al.
Published: (2026) -
Model Predictive Control is Almost Optimal for Restless Bandit
by: Gast, Nicolas, et al.
Published: (2024) -
Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits
by: Narasimha, Dheeraj, et al.
Published: (2025) -
Landscape of Policy Optimization for Finite Horizon MDPs with General State and Action
by: Chen, Xin, et al.
Published: (2024)