Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Harin, Oh, Min-hwan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improved Regret of Linear Ensemble Sampling
von: Lee, Harin, et al.
Veröffentlicht: (2024)
von: Lee, Harin, et al.
Veröffentlicht: (2024)
Infrequent Exploration in Linear Bandits
von: Lee, Harin, et al.
Veröffentlicht: (2025)
von: Lee, Harin, et al.
Veröffentlicht: (2025)
Minimax Optimal Reinforcement Learning with Quasi-Optimism
von: Lee, Harin, et al.
Veröffentlicht: (2025)
von: Lee, Harin, et al.
Veröffentlicht: (2025)
Lasso Bandit with Compatibility Condition on Optimal Arm
von: Lee, Harin, et al.
Veröffentlicht: (2024)
von: Lee, Harin, et al.
Veröffentlicht: (2024)
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
von: Lee, Joongkyu, et al.
Veröffentlicht: (2024)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2024)
Local Anti-Concentration Class: Logarithmic Regret for Greedy Linear Contextual Bandit
von: Kim, Seok-Jin, et al.
Veröffentlicht: (2024)
von: Kim, Seok-Jin, et al.
Veröffentlicht: (2024)
Improved Online Confidence Bounds for Multinomial Logistic Bandits
von: Lee, Joongkyu, et al.
Veröffentlicht: (2025)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2025)
Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent
von: Lee, Joongkyu, et al.
Veröffentlicht: (2026)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2026)
Blessings of Multiple Good Arms in Multi-Objective Linear Bandits
von: Ann, Heesang, et al.
Veröffentlicht: (2026)
von: Ann, Heesang, et al.
Veröffentlicht: (2026)
Combinatorial Reinforcement Learning with Preference Feedback
von: Lee, Joongkyu, et al.
Veröffentlicht: (2025)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2025)
Optimal and Practical Batched Linear Bandit Algorithm
von: Yu, Sanghoon, et al.
Veröffentlicht: (2025)
von: Yu, Sanghoon, et al.
Veröffentlicht: (2025)
Thompson Sampling for Multi-Objective Linear Contextual Bandit
von: Park, Somangchan, et al.
Veröffentlicht: (2025)
von: Park, Somangchan, et al.
Veröffentlicht: (2025)
Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates
von: Yu, Sanghoon, et al.
Veröffentlicht: (2026)
von: Yu, Sanghoon, et al.
Veröffentlicht: (2026)
Queueing Matching Bandits with Preference Feedback
von: Kim, Jung-hun, et al.
Veröffentlicht: (2024)
von: Kim, Jung-hun, et al.
Veröffentlicht: (2024)
Individual Regret in Cooperative Stochastic Multi-Armed Bandits
von: Barnea, Idan, et al.
Veröffentlicht: (2024)
von: Barnea, Idan, et al.
Veröffentlicht: (2024)
Follow-the-Perturbed-Leader for Decoupled Bandits: Best-of-Both-Worlds and Practicality
von: Kim, Chaiwon, et al.
Veröffentlicht: (2025)
von: Kim, Chaiwon, et al.
Veröffentlicht: (2025)
Follow-the-Perturbed-Leader with Fréchet-type Tail Distributions: Optimality in Adversarial Bandits and Best-of-Both-Worlds
von: Lee, Jongyeong, et al.
Veröffentlicht: (2024)
von: Lee, Jongyeong, et al.
Veröffentlicht: (2024)
Exploration via Feature Perturbation in Contextual Bandits
von: Yi, Seouh-won, et al.
Veröffentlicht: (2025)
von: Yi, Seouh-won, et al.
Veröffentlicht: (2025)
Stochastic Matching Bandits with Rare Optimization Updates
von: Kim, Jung-hun, et al.
Veröffentlicht: (2025)
von: Kim, Jung-hun, et al.
Veröffentlicht: (2025)
Model-Based Reinforcement Learning with Multinomial Logistic Function Approximation
von: Hwang, Taehyun, et al.
Veröffentlicht: (2022)
von: Hwang, Taehyun, et al.
Veröffentlicht: (2022)
Demystifying Linear MDPs and Novel Dynamics Aggregation Framework
von: Lee, Joongkyu, et al.
Veröffentlicht: (2024)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2024)
Graph-Dependent Regret Bounds in Multi-Armed Bandits with Interference
von: Jamshidi, Fateme, et al.
Veröffentlicht: (2025)
von: Jamshidi, Fateme, et al.
Veröffentlicht: (2025)
Collaborative Min-Max Regret in Grouped Multi-Armed Bandits
von: Blanchard, Moïse, et al.
Veröffentlicht: (2025)
von: Blanchard, Moïse, et al.
Veröffentlicht: (2025)
Peng's Q($λ$) for Conservative Value Estimation in Offline Reinforcement Learning
von: Kim, Byeongchan, et al.
Veröffentlicht: (2026)
von: Kim, Byeongchan, et al.
Veröffentlicht: (2026)
Revisiting Follow-the-Perturbed-Leader with Unbounded Perturbations in Bandit Problems
von: Lee, Jongyeong, et al.
Veröffentlicht: (2025)
von: Lee, Jongyeong, et al.
Veröffentlicht: (2025)
Tractable Multinomial Logit Contextual Bandits with Non-Linear Utilities
von: Hwang, Taehyun, et al.
Veröffentlicht: (2026)
von: Hwang, Taehyun, et al.
Veröffentlicht: (2026)
Adversarial Policy Optimization for Offline Preference-based Reinforcement Learning
von: Kang, Hyungkyu, et al.
Veröffentlicht: (2025)
von: Kang, Hyungkyu, et al.
Veröffentlicht: (2025)
Oracle-Efficient Combinatorial Semi-Bandits
von: Kim, Jung-hun, et al.
Veröffentlicht: (2025)
von: Kim, Jung-hun, et al.
Veröffentlicht: (2025)
Experimental Design for Semiparametric Bandits
von: Kim, Seok-Jin, et al.
Veröffentlicht: (2025)
von: Kim, Seok-Jin, et al.
Veröffentlicht: (2025)
Stochastic Multi-Objective Multi-Armed Bandits: Regret Definition and Algorithm
von: Davoodi, Mansoor, et al.
Veröffentlicht: (2025)
von: Davoodi, Mansoor, et al.
Veröffentlicht: (2025)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
von: Ji, Kaixuan, et al.
Veröffentlicht: (2026)
von: Ji, Kaixuan, et al.
Veröffentlicht: (2026)
On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits
von: Hou, Yunlong, et al.
Veröffentlicht: (2026)
von: Hou, Yunlong, et al.
Veröffentlicht: (2026)
Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation
von: Cho, Wooseong, et al.
Veröffentlicht: (2024)
von: Cho, Wooseong, et al.
Veröffentlicht: (2024)
Asymptotically and Minimax Optimal Regret Bounds for Multi-Armed Bandits with Abstention
von: Yang, Junwen, et al.
Veröffentlicht: (2024)
von: Yang, Junwen, et al.
Veröffentlicht: (2024)
Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning
von: Kang, Hyungkyu, et al.
Veröffentlicht: (2026)
von: Kang, Hyungkyu, et al.
Veröffentlicht: (2026)
Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
von: Lee, Joongkyu, et al.
Veröffentlicht: (2025)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2025)
Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification
von: Lee, Joongkyu, et al.
Veröffentlicht: (2026)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2026)
Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
von: Lee, Harin, et al.
Veröffentlicht: (2026)
von: Lee, Harin, et al.
Veröffentlicht: (2026)
Linear Bandits with Partially Observable Features
von: Kim, Wonyoung, et al.
Veröffentlicht: (2025)
von: Kim, Wonyoung, et al.
Veröffentlicht: (2025)
Understanding Memory-Regret Trade-Off for Streaming Stochastic Multi-Armed Bandits
von: He, Yuchen, et al.
Veröffentlicht: (2024)
von: He, Yuchen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improved Regret of Linear Ensemble Sampling
von: Lee, Harin, et al.
Veröffentlicht: (2024) -
Infrequent Exploration in Linear Bandits
von: Lee, Harin, et al.
Veröffentlicht: (2025) -
Minimax Optimal Reinforcement Learning with Quasi-Optimism
von: Lee, Harin, et al.
Veröffentlicht: (2025) -
Lasso Bandit with Compatibility Condition on Optimal Arm
von: Lee, Harin, et al.
Veröffentlicht: (2024) -
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
von: Lee, Joongkyu, et al.
Veröffentlicht: (2024)