Saved in:
| Main Authors: | Prevost, Adrien, Mathieu, Timothee, Maillard, Odalric-Ambrym |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.19098 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Leveraging priors on distribution functions for multi-arm bandits
by: Vashishtha, Sumit, et al.
Published: (2025)
by: Vashishtha, Sumit, et al.
Published: (2025)
How Hard is it to Confuse a World Model?
by: Radji, Waris, et al.
Published: (2025)
by: Radji, Waris, et al.
Published: (2025)
The regret lower bound for communicating Markov Decision Processes
by: Boone, Victor, et al.
Published: (2025)
by: Boone, Victor, et al.
Published: (2025)
The Confusing Instance Principle for Online Linear Quadratic Control
by: Radji, Waris, et al.
Published: (2025)
by: Radji, Waris, et al.
Published: (2025)
Hierarchical Subspaces of Policies for Continual Offline Reinforcement Learning
by: Kobanda, Anthony, et al.
Published: (2024)
by: Kobanda, Anthony, et al.
Published: (2024)
A Continual Offline Reinforcement Learning Benchmark for Navigation Tasks
by: Kobanda, Anthony, et al.
Published: (2025)
by: Kobanda, Anthony, et al.
Published: (2025)
How to Shrink Confidence Sets for Many Equivalent Discrete Distributions?
by: Maillard, Odalric-Ambrym, et al.
Published: (2024)
by: Maillard, Odalric-Ambrym, et al.
Published: (2024)
Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
by: Kobanda, Anthony, et al.
Published: (2025)
by: Kobanda, Anthony, et al.
Published: (2025)
Pliable rejection sampling
by: Erraqabi, Akram, et al.
Published: (2026)
by: Erraqabi, Akram, et al.
Published: (2026)
AdaStop: adaptive statistical testing for sound comparisons of Deep RL agents
by: Mathieu, Timothée, et al.
Published: (2023)
by: Mathieu, Timothée, et al.
Published: (2023)
Extended UCB Policies for Multi-armed Bandit Problems
by: Liu, Keqin, et al.
Published: (2011)
by: Liu, Keqin, et al.
Published: (2011)
Transfer Learning for Contextual Multi-armed Bandits
by: Cai, Changxiao, et al.
Published: (2022)
by: Cai, Changxiao, et al.
Published: (2022)
Optimal Batched Linear Bandits
by: Ren, Xuanfei, et al.
Published: (2024)
by: Ren, Xuanfei, et al.
Published: (2024)
A Simple and Optimal Policy Design with Safety against Heavy-Tailed Risk for Stochastic Bandits
by: Simchi-Levi, David, et al.
Published: (2022)
by: Simchi-Levi, David, et al.
Published: (2022)
Non-Asymptotic Analysis of Data Augmentation for Precision Matrix Estimation
by: Morisset, Lucas, et al.
Published: (2025)
by: Morisset, Lucas, et al.
Published: (2025)
Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms
by: Réveillard, William, et al.
Published: (2025)
by: Réveillard, William, et al.
Published: (2025)
Set-Valued Policy Learning
by: Fuentes-Vicente, Laura, et al.
Published: (2026)
by: Fuentes-Vicente, Laura, et al.
Published: (2026)
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
by: Praharaj, Samya, et al.
Published: (2025)
by: Praharaj, Samya, et al.
Published: (2025)
Statistical Complexity and Optimal Algorithms for Non-linear Ridge Bandits
by: Rajaraman, Nived, et al.
Published: (2023)
by: Rajaraman, Nived, et al.
Published: (2023)
Asymptotically Optimal Sequential Testing with Markovian Data
by: Sethi, Alhad, et al.
Published: (2026)
by: Sethi, Alhad, et al.
Published: (2026)
Minimax Rate-Optimal Algorithms for High-Dimensional Stochastic Linear Bandits
by: Liu, Jingyu, et al.
Published: (2025)
by: Liu, Jingyu, et al.
Published: (2025)
Towards Efficient and Optimal Covariance-Adaptive Algorithms for Combinatorial Semi-Bandits
by: Zhou, Julien, et al.
Published: (2024)
by: Zhou, Julien, et al.
Published: (2024)
Meta-Learning with Generalized Ridge Regression: High-dimensional Asymptotics, Optimality and Hyper-covariance Estimation
by: Jin, Yanhao, et al.
Published: (2024)
by: Jin, Yanhao, et al.
Published: (2024)
Provably Efficient Exploration in Reward Machines with Low Regret
by: Bourel, Hippolyte, et al.
Published: (2024)
by: Bourel, Hippolyte, et al.
Published: (2024)
Multitask Learning and Bandits via Robust Statistics
by: Xu, Kan, et al.
Published: (2021)
by: Xu, Kan, et al.
Published: (2021)
A Diffusion Analysis of Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)
by: Lattimore, Tor
Published: (2026)
Adaptive Lasso, Transfer Lasso, and Beyond: An Asymptotic Perspective
by: Takada, Masaaki, et al.
Published: (2023)
by: Takada, Masaaki, et al.
Published: (2023)
Source-Optimal Training is Transfer-Suboptimal
by: Hedges, C. Evans
Published: (2025)
by: Hedges, C. Evans
Published: (2025)
Optimal Transport under Group Fairness Constraints
by: Bleistein, Linus, et al.
Published: (2026)
by: Bleistein, Linus, et al.
Published: (2026)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025)
by: Zhao, Qingyue, et al.
Published: (2025)
Optimal Regret of Bernoulli Bandits under Global Differential Privacy
by: Azize, Achraf, et al.
Published: (2025)
by: Azize, Achraf, et al.
Published: (2025)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Laplace Transform Based Low-Complexity Learning of Continuous Markov Semigroups
by: Kostic, Vladimir R., et al.
Published: (2024)
by: Kostic, Vladimir R., et al.
Published: (2024)
Multi-Armed Bandits With Machine Learning-Generated Surrogate Rewards
by: Ji, Wenlong, et al.
Published: (2025)
by: Ji, Wenlong, et al.
Published: (2025)
Higher-Order Regularization Learning on Hypergraphs
by: Weihs, Adrien, et al.
Published: (2025)
by: Weihs, Adrien, et al.
Published: (2025)
Analysis of Semi-Supervised Learning on Hypergraphs
by: Weihs, Adrien, et al.
Published: (2025)
by: Weihs, Adrien, et al.
Published: (2025)
Aligning Embeddings and Geometric Random Graphs: Informational Results and Computational Approaches for the Procrustes-Wasserstein Problem
by: Even, Mathieu, et al.
Published: (2024)
by: Even, Mathieu, et al.
Published: (2024)
Regret Distribution in Stochastic Bandits: Optimal Trade-off between Expectation and Tail Risk
by: Simchi-Levi, David, et al.
Published: (2023)
by: Simchi-Levi, David, et al.
Published: (2023)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games
by: Ocello, Antonio, et al.
Published: (2025)
by: Ocello, Antonio, et al.
Published: (2025)
Similar Items
-
Leveraging priors on distribution functions for multi-arm bandits
by: Vashishtha, Sumit, et al.
Published: (2025) -
How Hard is it to Confuse a World Model?
by: Radji, Waris, et al.
Published: (2025) -
The regret lower bound for communicating Markov Decision Processes
by: Boone, Victor, et al.
Published: (2025) -
The Confusing Instance Principle for Online Linear Quadratic Control
by: Radji, Waris, et al.
Published: (2025) -
Hierarchical Subspaces of Policies for Continual Offline Reinforcement Learning
by: Kobanda, Anthony, et al.
Published: (2024)