DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Xiong, Guojun, Dinesha, Ujwal, Mukherjee, Debajoy, Li, Jian, Shakkottai, Srinivas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
by: Bobbili, Sarat Chandra, et al.
Published: (2025)
by: Bobbili, Sarat Chandra, et al.
Published: (2025)
Low-Complexity Algorithm for Restless Bandits with Imperfect Observations
by: Liu, Keqin, et al.
Published: (2021)
by: Liu, Keqin, et al.
Published: (2021)
Model Predictive Control is Almost Optimal for Restless Bandit
by: Gast, Nicolas, et al.
Published: (2024)
by: Gast, Nicolas, et al.
Published: (2024)
On the Linear Speedup of Personalized Federated Reinforcement Learning with Shared Representations
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
Model Predictive Control is almost Optimal for Heterogeneous Restless Multi-armed Bandits
by: Narasimha, Dheeraj, et al.
Published: (2025)
by: Narasimha, Dheeraj, et al.
Published: (2025)
Lagrangian Index Policy for Restless Bandits with Average Reward
by: Avrachenkov, Konstantin, et al.
Published: (2024)
by: Avrachenkov, Konstantin, et al.
Published: (2024)
Online Learning on Hidden-Convex Losses via Algorithmic Equivalence: Optimal Regret, Geometric Barrier, and Bandit Feedback
by: Barakat, Anas, et al.
Published: (2026)
by: Barakat, Anas, et al.
Published: (2026)
CONGO: Compressive Online Gradient Optimization
by: Carleton, Jeremy, et al.
Published: (2024)
by: Carleton, Jeremy, et al.
Published: (2024)
MAVIS: Multi-Objective Alignment via Inference-Time Value-Guided Selection
by: Carleton, Jeremy, et al.
Published: (2025)
by: Carleton, Jeremy, et al.
Published: (2025)
Distributed Online Bandit Nonconvex Optimization with One-Point Residual Feedback via Dynamic Regret
by: Hua, Youqing, et al.
Published: (2024)
by: Hua, Youqing, et al.
Published: (2024)
Preference-Optimized Pareto Set Learning for Blackbox Optimization
by: Haishan, Zhang, et al.
Published: (2024)
by: Haishan, Zhang, et al.
Published: (2024)
Online Newton Method for Bandit Convex Optimisation
by: Fokkema, Hidde, et al.
Published: (2024)
by: Fokkema, Hidde, et al.
Published: (2024)
FERERO: A Flexible Framework for Preference-Guided Multi-Objective Learning
by: Chen, Lisha, et al.
Published: (2024)
by: Chen, Lisha, et al.
Published: (2024)
Doubly Optimal No-Regret Online Learning in Strongly Monotone Games with Bandit Feedback
by: Ba, Wenjia, et al.
Published: (2021)
by: Ba, Wenjia, et al.
Published: (2021)
Combinatorial Causal Bandits without Graph Skeleton
by: Feng, Shi, et al.
Published: (2023)
by: Feng, Shi, et al.
Published: (2023)
PREFER: Personalized Review Summarization with Online Preference Learning
by: Roy, Millend, et al.
Published: (2026)
by: Roy, Millend, et al.
Published: (2026)
Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference Guidance
by: Chen, Lisha, et al.
Published: (2025)
by: Chen, Lisha, et al.
Published: (2025)
Follow The Approximate Sparse Leader for No-Regret Online Sparse Linear Approximation
by: Mukhopadhyay, Samrat, et al.
Published: (2025)
by: Mukhopadhyay, Samrat, et al.
Published: (2025)
A Regularized Online Newton Method for Stochastic Convex Bandits with Linear Vanishing Noise
by: Zhan, Jingxin, et al.
Published: (2025)
by: Zhan, Jingxin, et al.
Published: (2025)
Learning to Sparsify Stochastic Linear Bandits
by: Wang, Zhengmiao, et al.
Published: (2026)
by: Wang, Zhengmiao, et al.
Published: (2026)
Safe Online Convex Optimization with Multi-Point Feedback
by: Hutchinson, Spencer, et al.
Published: (2024)
by: Hutchinson, Spencer, et al.
Published: (2024)
Global and Preference-based Optimization with Mixed Variables using Piecewise Affine Surrogates
by: Zhu, Mengjia, et al.
Published: (2023)
by: Zhu, Mengjia, et al.
Published: (2023)
Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble
by: Wang, Hanyang, et al.
Published: (2025)
by: Wang, Hanyang, et al.
Published: (2025)
The Vizier Gaussian Process Bandit Algorithm
by: Song, Xingyou, et al.
Published: (2024)
by: Song, Xingyou, et al.
Published: (2024)
Safe and Efficient Online Convex Optimization with Linear Budget Constraints and Partial Feedback
by: Liu, Shanqi, et al.
Published: (2024)
by: Liu, Shanqi, et al.
Published: (2024)
Bandit Convex Optimisation
by: Lattimore, Tor
Published: (2024)
by: Lattimore, Tor
Published: (2024)
Adversarial Network Optimization under Bandit Feedback: Maximizing Utility in Non-Stationary Multi-Hop Networks
by: Dai, Yan, et al.
Published: (2024)
by: Dai, Yan, et al.
Published: (2024)
Restless Bandits with Average Reward: Breaking the Uniform Global Attractor Assumption
by: Hong, Yige, et al.
Published: (2023)
by: Hong, Yige, et al.
Published: (2023)
Feel-Good Thompson Sampling for Contextual Dueling Bandits
by: Li, Xuheng, et al.
Published: (2024)
by: Li, Xuheng, et al.
Published: (2024)
Inference of Utilities and Time Preference in Sequential Decision-Making
by: Cao, Haoyang, et al.
Published: (2024)
by: Cao, Haoyang, et al.
Published: (2024)
Learning to Cover: Online Learning and Optimization with Irreversible Decisions
by: Jacquillat, Alexandre, et al.
Published: (2024)
by: Jacquillat, Alexandre, et al.
Published: (2024)
Preference-aware compensation policies for crowdsourced on-demand services
by: Nouli, Georgina, et al.
Published: (2025)
by: Nouli, Georgina, et al.
Published: (2025)
Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits
by: Li, Xuheng, et al.
Published: (2025)
by: Li, Xuheng, et al.
Published: (2025)
Signature Approach for Contextual Bandits with Nonlinear and Path-dependent Rewards
by: Guo, Xin, et al.
Published: (2026)
by: Guo, Xin, et al.
Published: (2026)
Controlling Participation in Federated Learning with Feedback
by: Cummins, Michael, et al.
Published: (2024)
by: Cummins, Michael, et al.
Published: (2024)
Contextual Bandits with Budgeted Information Reveal
by: Gan, Kyra, et al.
Published: (2023)
by: Gan, Kyra, et al.
Published: (2023)
Decentralized Contextual Bandits with Network Adaptivity
by: Deng, Chuyun, et al.
Published: (2025)
by: Deng, Chuyun, et al.
Published: (2025)
The Safety-Privacy Tradeoff in Linear Bandits
by: Zibaie, Arghavan, et al.
Published: (2025)
by: Zibaie, Arghavan, et al.
Published: (2025)
Proximal Point Method for Online Saddle Point Problem
by: Meng, Qing-xin, et al.
Published: (2024)
by: Meng, Qing-xin, et al.
Published: (2024)
Similar Items
-
PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
by: Bobbili, Sarat Chandra, et al.
Published: (2025) -
Low-Complexity Algorithm for Restless Bandits with Imperfect Observations
by: Liu, Keqin, et al.
Published: (2021) -
Model Predictive Control is Almost Optimal for Restless Bandit
by: Gast, Nicolas, et al.
Published: (2024) -
On the Linear Speedup of Personalized Federated Reinforcement Learning with Shared Representations
by: Xiong, Guojun, et al.
Published: (2024) -
Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback
by: Xiong, Guojun, et al.
Published: (2024)