Lagrangian Index Policy for Restless Bandits with Average Reward
Fuente:
arXiv
Saved in:
| Main Authors: | Avrachenkov, Konstantin, Borkar, Vivek S., Shah, Pratik |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model Predictive Control is Almost Optimal for Restless Bandit
by: Gast, Nicolas, et al.
Published: (2024)
by: Gast, Nicolas, et al.
Published: (2024)
Restless Bandits with Average Reward: Breaking the Uniform Global Attractor Assumption
by: Hong, Yige, et al.
Published: (2023)
by: Hong, Yige, et al.
Published: (2023)
Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits
by: Narasimha, Dheeraj, et al.
Published: (2025)
by: Narasimha, Dheeraj, et al.
Published: (2025)
Achieving Exponential Asymptotic Optimality in Average-Reward Restless Bandits without Global Attractor Assumption
by: Hong, Yige, et al.
Published: (2024)
by: Hong, Yige, et al.
Published: (2024)
Tabular and Deep Learning for the Whittle Index
by: Relaño, Francisco Robledo, et al.
Published: (2024)
by: Relaño, Francisco Robledo, et al.
Published: (2024)
Linking PageRank, Time Reversal, and Policy Evaluation
by: Avrachenkov, Konstantin, et al.
Published: (2026)
by: Avrachenkov, Konstantin, et al.
Published: (2026)
Stability of Polling Systems for a Large Class of Markovian Switching Policies
by: Avrachenkov, Konstantin, et al.
Published: (2025)
by: Avrachenkov, Konstantin, et al.
Published: (2025)
Computing Stationary Distribution via Dirichlet-Energy Minimization by Coordinate Descent
by: Avrachenkov, Konstantin, et al.
Published: (2026)
by: Avrachenkov, Konstantin, et al.
Published: (2026)
Q-Learning under Finite Model Uncertainty
by: Sester, Julian, et al.
Published: (2024)
by: Sester, Julian, et al.
Published: (2024)
On the consistent reasoning paradox of intelligence and optimal trust in AI: The power of 'I don't know'
by: Bastounis, Alexander, et al.
Published: (2024)
by: Bastounis, Alexander, et al.
Published: (2024)
Neural Brownian Motion
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
Feature-aligned N-BEATS with Sinkhorn divergence
by: Lee, Joonhun, et al.
Published: (2023)
by: Lee, Joonhun, et al.
Published: (2023)
Robust $Q$-learning Algorithm for Markov Decision Processes under Wasserstein Uncertainty
by: Neufeld, Ariel, et al.
Published: (2022)
by: Neufeld, Ariel, et al.
Published: (2022)
Probabilistic Geometric Alignment via Bayesian Latent Transport for Domain-Adaptive Foundation Models
by: Aueawatthanaphisut, Aueaphum, et al.
Published: (2026)
by: Aueawatthanaphisut, Aueaphum, et al.
Published: (2026)
Efficient Risk-sensitive Planning via Entropic Risk Measures
by: Marthe, Alexandre, et al.
Published: (2025)
by: Marthe, Alexandre, et al.
Published: (2025)
Bandit Allocational Instability
by: Chen, Yilun, et al.
Published: (2026)
by: Chen, Yilun, et al.
Published: (2026)
Unichain and Aperiodicity are Sufficient for Asymptotic Optimality of Average-Reward Restless Bandits
by: Hong, Yige, et al.
Published: (2024)
by: Hong, Yige, et al.
Published: (2024)
Beyond Average Return in Markov Decision Processes
by: Marthe, Alexandre, et al.
Published: (2023)
by: Marthe, Alexandre, et al.
Published: (2023)
Derivative Estimation from Coarse, Irregular, Noisy Samples: An MLE-Spline Approach
by: Avrachenkov, Konstantin E., et al.
Published: (2025)
by: Avrachenkov, Konstantin E., et al.
Published: (2025)
Constrained Average-Reward Intermittently Observable MDPs
by: Avrachenkov, Konstantin, et al.
Published: (2025)
by: Avrachenkov, Konstantin, et al.
Published: (2025)
Representative Action Selection for Large Action Space Bandit Families
by: Zhou, Quan, et al.
Published: (2025)
by: Zhou, Quan, et al.
Published: (2025)
Representative Action Selection for Large Action Space: From Bandits to MDPs
by: Zhou, Quan, et al.
Published: (2025)
by: Zhou, Quan, et al.
Published: (2025)
Low-Complexity Algorithm for Restless Bandits with Imperfect Observations
by: Liu, Keqin, et al.
Published: (2021)
by: Liu, Keqin, et al.
Published: (2021)
Online Resource Allocation with Average Budget Constraints
by: Ao, Ruicheng, et al.
Published: (2024)
by: Ao, Ruicheng, et al.
Published: (2024)
The Vizier Gaussian Process Bandit Algorithm
by: Song, Xingyou, et al.
Published: (2024)
by: Song, Xingyou, et al.
Published: (2024)
Structure Matters: Dynamic Policy Gradient
by: Klein, Sara, et al.
Published: (2024)
by: Klein, Sara, et al.
Published: (2024)
Early Stopping in Contextual Bandits and Inferences
by: Cui, Zihan
Published: (2025)
by: Cui, Zihan
Published: (2025)
Power Constrained Nonstationary Bandits with Habituation and Recovery Dynamics
by: Li, Fengxu, et al.
Published: (2025)
by: Li, Fengxu, et al.
Published: (2025)
Pairwise independent correlation gap
by: Ramachandra, Arjun, et al.
Published: (2022)
by: Ramachandra, Arjun, et al.
Published: (2022)
DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
by: Mishchenko, Konstantin, et al.
Published: (2023)
by: Mishchenko, Konstantin, et al.
Published: (2023)
Wasserstein Formulation of Reinforcement Learning. An Optimal Transport Perspective on Policy Optimization
by: Dus, Mathias
Published: (2026)
by: Dus, Mathias
Published: (2026)
Asymptotically Optimal Policies for Weakly Coupled Markov Decision Processes
by: Goldsztajn, Diego, et al.
Published: (2024)
by: Goldsztajn, Diego, et al.
Published: (2024)
A dynamic view of some anomalous phenomena in SGD
by: Borkar, Vivek Shripad
Published: (2025)
by: Borkar, Vivek Shripad
Published: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
Accelerating RLHF Training with Reward Variance Increase
by: Yang, Zonglin, et al.
Published: (2025)
by: Yang, Zonglin, et al.
Published: (2025)
Gradient Descent Efficiency Index
by: Dhingra, Aviral
Published: (2024)
by: Dhingra, Aviral
Published: (2024)
The Gittins Index: A Design Principle for Decision-Making Under Uncertainty
by: Scully, Ziv, et al.
Published: (2025)
by: Scully, Ziv, et al.
Published: (2025)
The Implicit Curriculum: Learning Dynamics in RL with Verifiable Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Balans: Multi-Armed Bandits-based Adaptive Large Neighborhood Search for Mixed-Integer Programming Problem
by: Cai, Junyang, et al.
Published: (2024)
by: Cai, Junyang, et al.
Published: (2024)
Similar Items
-
Model Predictive Control is Almost Optimal for Restless Bandit
by: Gast, Nicolas, et al.
Published: (2024) -
Restless Bandits with Average Reward: Breaking the Uniform Global Attractor Assumption
by: Hong, Yige, et al.
Published: (2023) -
Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits
by: Narasimha, Dheeraj, et al.
Published: (2025) -
Achieving Exponential Asymptotic Optimality in Average-Reward Restless Bandits without Global Attractor Assumption
by: Hong, Yige, et al.
Published: (2024) -
Tabular and Deep Learning for the Whittle Index
by: Relaño, Francisco Robledo, et al.
Published: (2024)