Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Ito, Shinji, Jamieson, Kevin, Luo, Haipeng, Maiti, Arnab, Tsuchiya, Taira |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement Learning from Adversarial Preferences in Tabular MDPs
by: Tsuchiya, Taira, et al.
Published: (2025)
by: Tsuchiya, Taira, et al.
Published: (2025)
Adversarial Learning in Games with Bandit Feedback: Logarithmic Pure-Strategy Maximin Regret
by: Ito, Shinji, et al.
Published: (2026)
by: Ito, Shinji, et al.
Published: (2026)
Instance-Dependent Regret Bounds for Learning Two-Player Zero-Sum Games with Bandit Feedback
by: Ito, Shinji, et al.
Published: (2025)
by: Ito, Shinji, et al.
Published: (2025)
Scale-Invariant Fast Convergence in Games
by: Tsuchiya, Taira, et al.
Published: (2026)
by: Tsuchiya, Taira, et al.
Published: (2026)
Corrupted Learning Dynamics in Games
by: Tsuchiya, Taira, et al.
Published: (2024)
by: Tsuchiya, Taira, et al.
Published: (2024)
Fast Rates in Stochastic Online Convex Optimization by Exploiting the Curvature of Feasible Sets
by: Tsuchiya, Taira, et al.
Published: (2024)
by: Tsuchiya, Taira, et al.
Published: (2024)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
On the Power of Adaptivity for $\varepsilon$-Best Arm Identification in Linear Bandits
by: Maiti, Arnab, et al.
Published: (2026)
by: Maiti, Arnab, et al.
Published: (2026)
Exploration by Optimization with Hybrid Regularizers: Logarithmic Regret with Adversarial Robustness in Partial Monitoring
by: Tsuchiya, Taira, et al.
Published: (2024)
by: Tsuchiya, Taira, et al.
Published: (2024)
A Simple and Adaptive Learning Rate for FTRL in Online Learning with Minimax Regret of $Θ(T^{2/3})$ and its Application to Best-of-Both-Worlds
by: Tsuchiya, Taira, et al.
Published: (2024)
by: Tsuchiya, Taira, et al.
Published: (2024)
Efficient Near-Optimal Algorithm for Online Shortest Paths in Directed Acyclic Graphs with Bandit Feedback Against Adaptive Adversaries
by: Maiti, Arnab, et al.
Published: (2025)
by: Maiti, Arnab, et al.
Published: (2025)
Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds
by: Tsuchiya, Taira, et al.
Published: (2023)
by: Tsuchiya, Taira, et al.
Published: (2023)
Adaptive Learning Rate for Follow-the-Regularized-Leader: Competitive Analysis and Best-of-Both-Worlds
by: Ito, Shinji, et al.
Published: (2024)
by: Ito, Shinji, et al.
Published: (2024)
On the Limitations and Possibilities of Nash Regret Minimization in Zero-Sum Matrix Games under Noisy Feedback
by: Maiti, Arnab, et al.
Published: (2023)
by: Maiti, Arnab, et al.
Published: (2023)
Bandit and Delayed Feedback in Online Structured Prediction
by: Shibukawa, Yuki, et al.
Published: (2025)
by: Shibukawa, Yuki, et al.
Published: (2025)
Data- and Variance-dependent Regret Bounds for Online Tabular MDPs
by: Li, Mingyi, et al.
Published: (2026)
by: Li, Mingyi, et al.
Published: (2026)
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
by: Liu, Haolin, et al.
Published: (2024)
by: Liu, Haolin, et al.
Published: (2024)
Nearly Minimax Optimal Submodular Maximization with Bandit Feedback
by: Tajdini, Artin, et al.
Published: (2023)
by: Tajdini, Artin, et al.
Published: (2023)
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
Heavy-tailed Linear Bandits: Adversarial Robustness, Best-of-both-worlds, and Beyond
by: Zhao, Canzhe, et al.
Published: (2025)
by: Zhao, Canzhe, et al.
Published: (2025)
Learning to Incentivize in Repeated Principal-Agent Problems with Adversarial Agent Arrivals
by: Liu, Junyan, et al.
Published: (2025)
by: Liu, Junyan, et al.
Published: (2025)
Learning Kernel-Based MDPs from Episodic Preferential Feedback
by: Pavlovic, Nikola, et al.
Published: (2026)
by: Pavlovic, Nikola, et al.
Published: (2026)
Near-Optimal Regret in Adversarial Kernel Bandits
by: Zhang, Yu-Jie, et al.
Published: (2026)
by: Zhang, Yu-Jie, et al.
Published: (2026)
Online Control of Linear Systems under Unbounded Noise
by: Ito, Kaito, et al.
Published: (2024)
by: Ito, Kaito, et al.
Published: (2024)
Efficient Uncoupled Learning Dynamics with $\tilde{O}\!\left(T^{-1/4}\right)$ Last-Iterate Convergence in Bilinear Saddle-Point Problems over Convex Sets under Bandit Feedback
by: Maiti, Arnab, et al.
Published: (2026)
by: Maiti, Arnab, et al.
Published: (2026)
Efficient Contextual Bandits with Uninformed Feedback Graphs
by: Zhang, Mengxiao, et al.
Published: (2024)
by: Zhang, Mengxiao, et al.
Published: (2024)
Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
by: Kash, Ian A., et al.
Published: (2022)
by: Kash, Ian A., et al.
Published: (2022)
Learning Adversarial MDPs with Stochastic Hard Constraints
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
LC-Tsallis-INF: Generalized Best-of-Both-Worlds Linear Contextual Bandits
by: Kato, Masahiro, et al.
Published: (2024)
by: Kato, Masahiro, et al.
Published: (2024)
Influential Bandits: Pulling an Arm May Change the Environment
by: Sato, Ryoma, et al.
Published: (2025)
by: Sato, Ryoma, et al.
Published: (2025)
Bandit Max-Min Fair Allocation
by: Harada, Tsubasa, et al.
Published: (2025)
by: Harada, Tsubasa, et al.
Published: (2025)
Follow-the-Perturbed-Leader with Fréchet-type Tail Distributions: Optimality in Adversarial Bandits and Best-of-Both-Worlds
by: Lee, Jongyeong, et al.
Published: (2024)
by: Lee, Jongyeong, et al.
Published: (2024)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
by: Schlisselberg, Ofir, et al.
Published: (2026)
by: Schlisselberg, Ofir, et al.
Published: (2026)
Online Structured Prediction with Fenchel--Young Losses and Improved Surrogate Regret for Online Multiclass Classification with Logistic Loss
by: Sakaue, Shinsaku, et al.
Published: (2024)
by: Sakaue, Shinsaku, et al.
Published: (2024)
Bandits with Stochastic Experts: Constant Regret, Empirical Experts and Episodes
by: Sharma, Nihal, et al.
Published: (2021)
by: Sharma, Nihal, et al.
Published: (2021)
Revisiting Online Learning Approach to Inverse Linear Optimization: A Fenchel$-$Young Loss Perspective and Gap-Dependent Regret Analysis
by: Sakaue, Shinsaku, et al.
Published: (2025)
by: Sakaue, Shinsaku, et al.
Published: (2025)
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
by: Hait, Soumita, et al.
Published: (2026)
by: Hait, Soumita, et al.
Published: (2026)
Best-of-Both-Worlds Algorithms for Linear Contextual Bandits
by: Kuroki, Yuko, et al.
Published: (2023)
by: Kuroki, Yuko, et al.
Published: (2023)
An Axiomatic Approach to Loss Aggregation and an Adapted Aggregating Algorithm
by: Pacheco, Armando J. Cabrera, et al.
Published: (2024)
by: Pacheco, Armando J. Cabrera, et al.
Published: (2024)
Similar Items
-
Reinforcement Learning from Adversarial Preferences in Tabular MDPs
by: Tsuchiya, Taira, et al.
Published: (2025) -
Adversarial Learning in Games with Bandit Feedback: Logarithmic Pure-Strategy Maximin Regret
by: Ito, Shinji, et al.
Published: (2026) -
Instance-Dependent Regret Bounds for Learning Two-Player Zero-Sum Games with Bandit Feedback
by: Ito, Shinji, et al.
Published: (2025) -
Scale-Invariant Fast Convergence in Games
by: Tsuchiya, Taira, et al.
Published: (2026) -
Corrupted Learning Dynamics in Games
by: Tsuchiya, Taira, et al.
Published: (2024)