Saved in:
| Main Authors: | Pavlovic, Nikola, Vakili, Sattar, Zhao, Qing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.23650 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open Problem: Order Optimal Regret Bounds for Kernel-Based Reinforcement Learning
by: Vakili, Sattar
Published: (2024)
by: Vakili, Sattar
Published: (2024)
A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
by: Lazzaro, Joseph, et al.
Published: (2026)
by: Lazzaro, Joseph, et al.
Published: (2026)
Kernel-Based Function Approximation for Average Reward Reinforcement Learning: An Optimist No-Regret Algorithm
by: Vakili, Sattar, et al.
Published: (2024)
by: Vakili, Sattar, et al.
Published: (2024)
Kernelized Reinforcement Learning with Order Optimal Regret Bounds
by: Vakili, Sattar, et al.
Published: (2023)
by: Vakili, Sattar, et al.
Published: (2023)
Differentially Private Kernelized Contextual Bandits
by: Pavlovic, Nikola, et al.
Published: (2025)
by: Pavlovic, Nikola, et al.
Published: (2025)
Differential Privacy in Kernelized Contextual Bandits via Random Projections
by: Pavlovic, Nikola, et al.
Published: (2025)
by: Pavlovic, Nikola, et al.
Published: (2025)
Near-Optimal Sample Complexity in Reward-Free Kernel-Based Reinforcement Learning
by: Kayal, Aya, et al.
Published: (2025)
by: Kayal, Aya, et al.
Published: (2025)
Random Exploration in Bayesian Optimization: Order-Optimal Regret and Computational Efficiency
by: Salgia, Sudeep, et al.
Published: (2023)
by: Salgia, Sudeep, et al.
Published: (2023)
Order-Optimal Regret in Distributed Kernel Bandits using Uniform Sampling with Shared Randomness
by: Pavlovic, Nikola, et al.
Published: (2024)
by: Pavlovic, Nikola, et al.
Published: (2024)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
by: Ito, Shinji, et al.
Published: (2025)
by: Ito, Shinji, et al.
Published: (2025)
Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
by: Kayal, Aya, et al.
Published: (2025)
by: Kayal, Aya, et al.
Published: (2025)
Characterizing the Accuracy-Communication-Privacy Trade-off in Distributed Stochastic Convex Optimization
by: Salgia, Sudeep, et al.
Published: (2025)
by: Salgia, Sudeep, et al.
Published: (2025)
Reinforcement Learning Using known Invariances
by: Cioba, Alexandru, et al.
Published: (2025)
by: Cioba, Alexandru, et al.
Published: (2025)
No-Regret Thompson Sampling for Finite-Horizon Markov Decision Processes with Gaussian Processes
by: Bayrooti, Jasmine, et al.
Published: (2025)
by: Bayrooti, Jasmine, et al.
Published: (2025)
On the Convergence of Monte Carlo UCB for Random-Length Episodic MDPs
by: Dong, Zixuan, et al.
Published: (2022)
by: Dong, Zixuan, et al.
Published: (2022)
Reinforcement Learning from Multi-level and Episodic Human Feedback
by: Elahi, Muhammad Qasim, et al.
Published: (2025)
by: Elahi, Muhammad Qasim, et al.
Published: (2025)
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
by: Kitamura, Toshinori, et al.
Published: (2025)
by: Kitamura, Toshinori, et al.
Published: (2025)
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
by: Li, Long-Fei, et al.
Published: (2024)
by: Li, Long-Fei, et al.
Published: (2024)
Many Needles in a Haystack: Active Hit Discovery for Perturbation Experiments
by: Rubbi, Andrea, et al.
Published: (2026)
by: Rubbi, Andrea, et al.
Published: (2026)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
by: Liu, Haolin, et al.
Published: (2024)
by: Liu, Haolin, et al.
Published: (2024)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
by: Wang, Kaixin, et al.
Published: (2023)
by: Wang, Kaixin, et al.
Published: (2023)
Robust Parameter Learning for Uncertain MDPs
by: Schnitzer, Yannik, et al.
Published: (2026)
by: Schnitzer, Yannik, et al.
Published: (2026)
Truly No-Regret Learning in Constrained MDPs
by: Müller, Adrian, et al.
Published: (2024)
by: Müller, Adrian, et al.
Published: (2024)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
Preferential Normalizing Flows
by: Mikkola, Petrus, et al.
Published: (2024)
by: Mikkola, Petrus, et al.
Published: (2024)
Anchor-Based Heteroscedastic Noise for Preferential Bayesian Optimization
by: Sinaga, Marshal Arijona, et al.
Published: (2024)
by: Sinaga, Marshal Arijona, et al.
Published: (2024)
Reinforcement Learning from Adversarial Preferences in Tabular MDPs
by: Tsuchiya, Taira, et al.
Published: (2025)
by: Tsuchiya, Taira, et al.
Published: (2025)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Learning Adversarial MDPs with Stochastic Hard Constraints
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
Local Preferential Bayesian Optimization
by: Menn, Johanna, et al.
Published: (2026)
by: Menn, Johanna, et al.
Published: (2026)
Principled Preferential Bayesian Optimization
by: Xu, Wenjie, et al.
Published: (2024)
by: Xu, Wenjie, et al.
Published: (2024)
Consecutive Preferential Bayesian Optimization
by: Erarslan, Aras, et al.
Published: (2025)
by: Erarslan, Aras, et al.
Published: (2025)
Nonparametric Kernel Clustering with Bandit Feedback
by: Thuot, Victor, et al.
Published: (2026)
by: Thuot, Victor, et al.
Published: (2026)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
A Context-Aware Temporal Modeling through Unified Multi-Scale Temporal Encoding and Hierarchical Sequence Learning for Single-Channel EEG Sleep Staging
by: Vakili, Amirali, et al.
Published: (2025)
by: Vakili, Amirali, et al.
Published: (2025)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
by: Schlisselberg, Ofir, et al.
Published: (2026)
by: Schlisselberg, Ofir, et al.
Published: (2026)
Online Episodic Convex Reinforcement Learning
by: Moreno, Bianca Marin, et al.
Published: (2025)
by: Moreno, Bianca Marin, et al.
Published: (2025)
Koopman Learning with Episodic Memory
by: Redman, William T., et al.
Published: (2023)
by: Redman, William T., et al.
Published: (2023)
Efficient Solution and Learning of Robust Factored MDPs
by: Schnitzer, Yannik, et al.
Published: (2025)
by: Schnitzer, Yannik, et al.
Published: (2025)
Similar Items
-
Open Problem: Order Optimal Regret Bounds for Kernel-Based Reinforcement Learning
by: Vakili, Sattar
Published: (2024) -
A Finite Time Analysis of Thompson Sampling for Bayesian Optimization with Preferential Feedback
by: Lazzaro, Joseph, et al.
Published: (2026) -
Kernel-Based Function Approximation for Average Reward Reinforcement Learning: An Optimist No-Regret Algorithm
by: Vakili, Sattar, et al.
Published: (2024) -
Kernelized Reinforcement Learning with Order Optimal Regret Bounds
by: Vakili, Sattar, et al.
Published: (2023) -
Differentially Private Kernelized Contextual Bandits
by: Pavlovic, Nikola, et al.
Published: (2025)