Whittle Index Learning Algorithms for Restless Bandits with Constant Stepsizes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mittal, Vishesh, Meshram, Rahul, Prakash, Surya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Faster Q-Learning Algorithms for Restless Bandits
von: Kakarapalli, Parvish, et al.
Veröffentlicht: (2024)
von: Kakarapalli, Parvish, et al.
Veröffentlicht: (2024)
Lagrangian Relaxation for Multi-Action Partially Observable Restless Bandits: Heuristic Policies and Indexability
von: Meshram, Rahul, et al.
Veröffentlicht: (2025)
von: Meshram, Rahul, et al.
Veröffentlicht: (2025)
Risk-Aware Decision Making in Restless Bandits: Theory and Algorithms for Planning and Learning
von: Akbarzadeh, Nima, et al.
Veröffentlicht: (2024)
von: Akbarzadeh, Nima, et al.
Veröffentlicht: (2024)
Hierarchical Decentralized Stochastic Control for Cyber-Physical Systems
von: Kaza, Kesav, et al.
Veröffentlicht: (2025)
von: Kaza, Kesav, et al.
Veröffentlicht: (2025)
Restless Bandit Problem with Rewards Generated by a Linear Gaussian Dynamical System
von: Gornet, Jonathan, et al.
Veröffentlicht: (2024)
von: Gornet, Jonathan, et al.
Veröffentlicht: (2024)
Two-Timescale Linear Stochastic Approximation: Constant Stepsizes Go a Long Way
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Low-Complexity Algorithm for Restless Bandits with Imperfect Observations
von: Liu, Keqin, et al.
Veröffentlicht: (2021)
von: Liu, Keqin, et al.
Veröffentlicht: (2021)
The Bandit Whisperer: Communication Learning for Restless Bandits
von: Zhao, Yunfan, et al.
Veröffentlicht: (2024)
von: Zhao, Yunfan, et al.
Veröffentlicht: (2024)
Lagrangian Index Policy for Restless Bandits with Average Reward
von: Avrachenkov, Konstantin, et al.
Veröffentlicht: (2024)
von: Avrachenkov, Konstantin, et al.
Veröffentlicht: (2024)
Bandit Algorithms for Deep Brain Stimulation
von: Gupta, Arkaprava, et al.
Veröffentlicht: (2026)
von: Gupta, Arkaprava, et al.
Veröffentlicht: (2026)
Online Learning of Whittle Indices for Restless Bandits with Non-Stationary Transition Kernels
von: Shisher, Md Kamran Chowdhury, et al.
Veröffentlicht: (2025)
von: Shisher, Md Kamran Chowdhury, et al.
Veröffentlicht: (2025)
Learning to Sparsify Stochastic Linear Bandits
von: Wang, Zhengmiao, et al.
Veröffentlicht: (2026)
von: Wang, Zhengmiao, et al.
Veröffentlicht: (2026)
DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
von: Xiong, Guojun, et al.
Veröffentlicht: (2024)
von: Xiong, Guojun, et al.
Veröffentlicht: (2024)
Decentralized Upper Confidence Bound Algorithms for Homogeneous Multi-Agent Multi-Armed Bandits
von: Zhu, Jingxuan, et al.
Veröffentlicht: (2021)
von: Zhu, Jingxuan, et al.
Veröffentlicht: (2021)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
Bias and Extrapolation in Markovian Linear Stochastic Approximation with Constant Stepsizes
von: Huo, Dongyan, et al.
Veröffentlicht: (2022)
von: Huo, Dongyan, et al.
Veröffentlicht: (2022)
Finite-Horizon Single-Pull Restless Bandits: An Efficient Index Policy For Scarce Resource Allocation
von: Xiong, Guojun, et al.
Veröffentlicht: (2025)
von: Xiong, Guojun, et al.
Veröffentlicht: (2025)
ContextWIN: Whittle Index Based Mixture-of-Experts Neural Model For Restless Bandits Via Deep RL
von: Guo, Zhanqiu, et al.
Veröffentlicht: (2024)
von: Guo, Zhanqiu, et al.
Veröffentlicht: (2024)
Why Most Optimism Bandit Algorithms Have the Same Regret Analysis: A Simple Unifying Theorem
von: Krishnamurthy, Vikram
Veröffentlicht: (2025)
von: Krishnamurthy, Vikram
Veröffentlicht: (2025)
Model Predictive Control is Almost Optimal for Restless Bandit
von: Gast, Nicolas, et al.
Veröffentlicht: (2024)
von: Gast, Nicolas, et al.
Veröffentlicht: (2024)
The Collusion of Memory and Nonlinearity in Stochastic Approximation With Constant Stepsize
von: Huo, Dongyan, et al.
Veröffentlicht: (2024)
von: Huo, Dongyan, et al.
Veröffentlicht: (2024)
Byzantine-Resilient Decentralized Multi-Armed Bandits
von: Zhu, Jingxuan, et al.
Veröffentlicht: (2023)
von: Zhu, Jingxuan, et al.
Veröffentlicht: (2023)
Fairness for Workers Who Pull the Arms: An Index Based Policy for Allocation of Restless Bandit Tasks
von: Biswas, Arpita, et al.
Veröffentlicht: (2023)
von: Biswas, Arpita, et al.
Veröffentlicht: (2023)
Explore-then-Commit for Nonstationary Linear Bandits with Latent Dynamics
von: Choi, Sunmook, et al.
Veröffentlicht: (2025)
von: Choi, Sunmook, et al.
Veröffentlicht: (2025)
Multi-Agent Stage-wise Conservative Linear Bandits
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2025)
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2025)
HPC Application Parameter Autotuning on Edge Devices: A Bandit Learning Approach
von: Hossain, Abrar, et al.
Veröffentlicht: (2025)
von: Hossain, Abrar, et al.
Veröffentlicht: (2025)
Cost-Ordered Feasibility for Multi-Armed Bandits with Cost Subsidy
von: Juneja, Ishank, et al.
Veröffentlicht: (2026)
von: Juneja, Ishank, et al.
Veröffentlicht: (2026)
Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits
von: Narasimha, Dheeraj, et al.
Veröffentlicht: (2025)
von: Narasimha, Dheeraj, et al.
Veröffentlicht: (2025)
Multi-agent Multi-armed Bandits with Minimum Reward Guarantee Fairness
von: Manupriya, Piyushi, et al.
Veröffentlicht: (2025)
von: Manupriya, Piyushi, et al.
Veröffentlicht: (2025)
Finite-Time Guarantees for Multi-Agent Combinatorial Bandits with Nonstationary Rewards
von: Adams, Katherine B., et al.
Veröffentlicht: (2025)
von: Adams, Katherine B., et al.
Veröffentlicht: (2025)
Tangential Randomization in Linear Bandits (TRAiL): Guaranteed Inference and Regret Bounds
von: Güçlü, Arda, et al.
Veröffentlicht: (2024)
von: Güçlü, Arda, et al.
Veröffentlicht: (2024)
Differentially Private High Dimensional Bandits
von: Shukla, Apurv
Veröffentlicht: (2024)
von: Shukla, Apurv
Veröffentlicht: (2024)
Multi-User mmWave Beam and Rate Adaptation via Combinatorial Satisficing Bandits
von: Özyıldırım, Emre, et al.
Veröffentlicht: (2026)
von: Özyıldırım, Emre, et al.
Veröffentlicht: (2026)
An LP-based Sampling Policy for Multi-Armed Bandits with Side-Observations and Stochastic Availability
von: Soni, Ashutosh, et al.
Veröffentlicht: (2026)
von: Soni, Ashutosh, et al.
Veröffentlicht: (2026)
Adaptive Scheduling: A Reinforcement Learning Whittle Index Approach for Wireless Sensor Networks
von: Jonah, Sokipriala, et al.
Veröffentlicht: (2026)
von: Jonah, Sokipriala, et al.
Veröffentlicht: (2026)
ECLipsE: Efficient Compositional Lipschitz Constant Estimation for Deep Neural Networks
von: Xu, Yuezhu, et al.
Veröffentlicht: (2024)
von: Xu, Yuezhu, et al.
Veröffentlicht: (2024)
An Exploration-free Method for a Linear Stochastic Bandit Driven by a Linear Gaussian Dynamical System
von: Gornet, Jonathan, et al.
Veröffentlicht: (2025)
von: Gornet, Jonathan, et al.
Veröffentlicht: (2025)
Spatio-Temporal Graph Neural Networks for Dairy Farm Sustainability Forecasting and Counterfactual Policy Analysis
von: Jayakumar, Surya, et al.
Veröffentlicht: (2025)
von: Jayakumar, Surya, et al.
Veröffentlicht: (2025)
A Safe Reinforcement Learning Algorithm for Supervisory Control of Power Plants
von: Sun, Yixuan, et al.
Veröffentlicht: (2024)
von: Sun, Yixuan, et al.
Veröffentlicht: (2024)
Smoothed Online Optimization for Target Tracking: Robust and Learning-Augmented Algorithms
von: Zeynali, Ali, et al.
Veröffentlicht: (2025)
von: Zeynali, Ali, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Faster Q-Learning Algorithms for Restless Bandits
von: Kakarapalli, Parvish, et al.
Veröffentlicht: (2024) -
Lagrangian Relaxation for Multi-Action Partially Observable Restless Bandits: Heuristic Policies and Indexability
von: Meshram, Rahul, et al.
Veröffentlicht: (2025) -
Risk-Aware Decision Making in Restless Bandits: Theory and Algorithms for Planning and Learning
von: Akbarzadeh, Nima, et al.
Veröffentlicht: (2024) -
Hierarchical Decentralized Stochastic Control for Cyber-Physical Systems
von: Kaza, Kesav, et al.
Veröffentlicht: (2025) -
Restless Bandit Problem with Rewards Generated by a Linear Gaussian Dynamical System
von: Gornet, Jonathan, et al.
Veröffentlicht: (2024)