Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | John, Philips George, Bhattacharyya, Arnab, Maniu, Silviu, Myrisiotis, Dimitrios, Wu, Zhenan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distribution Learning Meets Graph Structure Sampling
by: Bhattacharyya, Arnab, et al.
Published: (2024)
by: Bhattacharyya, Arnab, et al.
Published: (2024)
Learning multivariate Gaussians with imperfect advice
by: Bhattacharyya, Arnab, et al.
Published: (2024)
by: Bhattacharyya, Arnab, et al.
Published: (2024)
Towards Optimal Differentially Private Regret Bounds in Linear MDPs
by: Sahu, Sharan
Published: (2025)
by: Sahu, Sharan
Published: (2025)
Total Variation Distance Meets Probabilistic Inference
by: Bhattacharyya, Arnab, et al.
Published: (2023)
by: Bhattacharyya, Arnab, et al.
Published: (2023)
Testing Sparse Functions over the Reals
by: Arora, Vipul, et al.
Published: (2026)
by: Arora, Vipul, et al.
Published: (2026)
A Single-Sample Polylogarithmic Regret Bound for Nonstationary Online Linear Programming
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Finite and Corruption-Robust Regret Bounds in Online Inverse Linear Optimization under M-Convex Action Sets
by: Oki, Taihei, et al.
Published: (2026)
by: Oki, Taihei, et al.
Published: (2026)
Online bipartite matching with imperfect advice
by: Choo, Davin, et al.
Published: (2024)
by: Choo, Davin, et al.
Published: (2024)
Computational Explorations of Total Variation Distance
by: Bhattacharyya, Arnab, et al.
Published: (2024)
by: Bhattacharyya, Arnab, et al.
Published: (2024)
Algorithms and Hardness for Estimating Statistical Similarity
by: Bhattacharyya, Arnab, et al.
Published: (2025)
by: Bhattacharyya, Arnab, et al.
Published: (2025)
Approximating the Total Variation Distance between Gaussians
by: Bhattacharyya, Arnab, et al.
Published: (2025)
by: Bhattacharyya, Arnab, et al.
Published: (2025)
Online Algorithms for Repeated Optimal Stopping: Balancing Baseline Guarantees and Regret
by: Harada, Tsubasa, et al.
Published: (2025)
by: Harada, Tsubasa, et al.
Published: (2025)
Improved Regret in Stochastic Decision-Theoretic Online Learning under Differential Privacy
by: Wu, Ruihan, et al.
Published: (2025)
by: Wu, Ruihan, et al.
Published: (2025)
Outlier Robust Multivariate Polynomial Regression
by: Arora, Vipul, et al.
Published: (2024)
by: Arora, Vipul, et al.
Published: (2024)
Near-Optimal Regret for Efficient Stochastic Combinatorial Semi-Bandits
by: Ye, Zichun, et al.
Published: (2025)
by: Ye, Zichun, et al.
Published: (2025)
$O(\sqrt{T})$ Static Regret and Instance Dependent Constraint Violation for Constrained Online Convex Optimization
by: Vaze, Rahul, et al.
Published: (2025)
by: Vaze, Rahul, et al.
Published: (2025)
Near-optimal Swap Regret Minimization for Convex Losses
by: Hu, Lunjia, et al.
Published: (2026)
by: Hu, Lunjia, et al.
Published: (2026)
Online Bilevel Optimization: Regret Analysis of Online Alternating Gradient Methods
by: Tarzanagh, Davoud Ataee, et al.
Published: (2022)
by: Tarzanagh, Davoud Ataee, et al.
Published: (2022)
LevAttention: Time, Space, and Streaming Efficient Algorithm for Heavy Attentions
by: Kannan, Ravindran, et al.
Published: (2024)
by: Kannan, Ravindran, et al.
Published: (2024)
Learning bounded-degree polytrees with known skeleton
by: Choo, Davin, et al.
Published: (2023)
by: Choo, Davin, et al.
Published: (2023)
Stochastic Multi-Objective Multi-Armed Bandits: Regret Definition and Algorithm
by: Davoodi, Mansoor, et al.
Published: (2025)
by: Davoodi, Mansoor, et al.
Published: (2025)
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
Improved and Oracle-Efficient Online $\ell_1$-Multicalibration
by: Ghuge, Rohan, et al.
Published: (2025)
by: Ghuge, Rohan, et al.
Published: (2025)
Attribute-Efficient PAC Learning of Low-Degree Polynomial Threshold Functions with Nasty Noise
by: Zeng, Shiwei, et al.
Published: (2023)
by: Zeng, Shiwei, et al.
Published: (2023)
Understanding Memory-Regret Trade-Off for Streaming Stochastic Multi-Armed Bandits
by: He, Yuchen, et al.
Published: (2024)
by: He, Yuchen, et al.
Published: (2024)
From Average Sensitivity to Small-Loss Regret Bounds under Random-Order Model
by: Sakaue, Shinsaku, et al.
Published: (2026)
by: Sakaue, Shinsaku, et al.
Published: (2026)
Tradeoffs between Mistakes and ERM Oracle Calls in Online and Transductive Online Learning
by: Attias, Idan, et al.
Published: (2025)
by: Attias, Idan, et al.
Published: (2025)
Online Learning in the Random Order Model
by: Bernasconi, Martino, et al.
Published: (2025)
by: Bernasconi, Martino, et al.
Published: (2025)
Transductive and Learning-Augmented Online Regression
by: Raman, Vinod, et al.
Published: (2025)
by: Raman, Vinod, et al.
Published: (2025)
Learning on the Edge: Online Learning with Stochastic Feedback Graphs
by: Esposito, Emmanuel, et al.
Published: (2022)
by: Esposito, Emmanuel, et al.
Published: (2022)
Simple Opinion Dynamics for No-Regret Learning
by: Lazarsfeld, John, et al.
Published: (2023)
by: Lazarsfeld, John, et al.
Published: (2023)
Learning-Augmented Algorithms for $k$-median via Online Learning
by: Hebbar, Anish, et al.
Published: (2026)
by: Hebbar, Anish, et al.
Published: (2026)
Parsimonious Learning-Augmented Online Metric Matching
by: Shin, Yongho, et al.
Published: (2026)
by: Shin, Yongho, et al.
Published: (2026)
Learning-Augmented Online Scheduling with Parsimonious Preemption
by: Blue, Mugen, et al.
Published: (2026)
by: Blue, Mugen, et al.
Published: (2026)
Efficient Algorithms for Verifying Kruskal Rank in Sparse Linear Regression and Related Applications
by: Zhou, Fengqin
Published: (2025)
by: Zhou, Fengqin
Published: (2025)
Learning Low Degree Hypergraphs
by: Balkanski, Eric, et al.
Published: (2022)
by: Balkanski, Eric, et al.
Published: (2022)
Online Learning with Limited Information in the Sliding Window Model
by: Braverman, Vladimir, et al.
Published: (2026)
by: Braverman, Vladimir, et al.
Published: (2026)
Online Linear Programming with Replenishment
by: Chen, Yuze, et al.
Published: (2026)
by: Chen, Yuze, et al.
Published: (2026)
No-Regret M${}^{\natural}$-Concave Function Maximization: Stochastic Bandit Algorithms and Hardness of Adversarial Full-Information Setting
by: Oki, Taihei, et al.
Published: (2024)
by: Oki, Taihei, et al.
Published: (2024)
Tight Gap-Dependent Memory-Regret Trade-Off for Single-Pass Streaming Stochastic Multi-Armed Bandits
by: Ye, Zichun, et al.
Published: (2025)
by: Ye, Zichun, et al.
Published: (2025)
Similar Items
-
Distribution Learning Meets Graph Structure Sampling
by: Bhattacharyya, Arnab, et al.
Published: (2024) -
Learning multivariate Gaussians with imperfect advice
by: Bhattacharyya, Arnab, et al.
Published: (2024) -
Towards Optimal Differentially Private Regret Bounds in Linear MDPs
by: Sahu, Sharan
Published: (2025) -
Total Variation Distance Meets Probabilistic Inference
by: Bhattacharyya, Arnab, et al.
Published: (2023) -
Testing Sparse Functions over the Reals
by: Arora, Vipul, et al.
Published: (2026)