Saved in:
| Main Authors: | Jia, Zeyu, Qian, Jian, Rakhlin, Alexander, Wei, Chen-Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2410.12713 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Minimax Regret of Sequential Probability Assignment via Square-Root Entropy
by: Jia, Zeyu, et al.
Published: (2025)
by: Jia, Zeyu, et al.
Published: (2025)
Self-Normalized Martingales and Uniform Regret Bounds for Linear Regression
by: Chen, Fan, et al.
Published: (2026)
by: Chen, Fan, et al.
Published: (2026)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
by: He, Jiafan, et al.
Published: (2025)
by: He, Jiafan, et al.
Published: (2025)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
Achieving Optimal Static and Dynamic Regret Simultaneously in Bandits with Deterministic Losses
by: Qian, Jian, et al.
Published: (2026)
by: Qian, Jian, et al.
Published: (2026)
The Sample Complexity of Multiclass and Sparse Contextual Bandits
by: Erez, Liad, et al.
Published: (2026)
by: Erez, Liad, et al.
Published: (2026)
Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data
by: Jia, Zeyu, et al.
Published: (2024)
by: Jia, Zeyu, et al.
Published: (2024)
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
by: Jia, Zeyu, et al.
Published: (2025)
by: Jia, Zeyu, et al.
Published: (2025)
A Gapped Scale-Sensitive Dimension and Lower Bounds for Offset Rademacher Complexity
by: Jia, Zeyu, et al.
Published: (2025)
by: Jia, Zeyu, et al.
Published: (2025)
On the Variance, Admissibility, and Stability of Empirical Risk Minimization
by: Kur, Gil, et al.
Published: (2023)
by: Kur, Gil, et al.
Published: (2023)
Batch Ensemble for Variance Dependent Regret in Stochastic Bandits
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Fast Best-in-Class Regret for Contextual Bandits
by: Girard, Samuel, et al.
Published: (2025)
by: Girard, Samuel, et al.
Published: (2025)
Optimal Regret for Policy Optimization in Contextual Bandits
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
Refined Risk Bounds for Unbounded Losses via Transductive Priors
by: Qian, Jian, et al.
Published: (2024)
by: Qian, Jian, et al.
Published: (2024)
Queue Length Regret Bounds for Contextual Queueing Bandits
by: Bae, Seoungbin, et al.
Published: (2026)
by: Bae, Seoungbin, et al.
Published: (2026)
Assouad, Fano, and Le Cam with Interaction: A Unifying Lower Bound Framework and Characterization for Bandit Learnability
by: Chen, Fan, et al.
Published: (2024)
by: Chen, Fan, et al.
Published: (2024)
On the Optimal Regret of Locally Private Linear Contextual Bandit
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
Active Context Selection Improves Simple Regret in Contextual Bandits
by: Shahverdikondori, Mohammad, et al.
Published: (2026)
by: Shahverdikondori, Mohammad, et al.
Published: (2026)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
by: Yuan, Yurun, et al.
Published: (2025)
by: Yuan, Yurun, et al.
Published: (2025)
Variance-Aware Linear UCB with Deep Representation for Neural Contextual Bandits
by: Bui, Ha Manh, et al.
Published: (2024)
by: Bui, Ha Manh, et al.
Published: (2024)
Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits
by: Li, Xuheng, et al.
Published: (2025)
by: Li, Xuheng, et al.
Published: (2025)
Local Anti-Concentration Class: Logarithmic Regret for Greedy Linear Contextual Bandit
by: Kim, Seok-Jin, et al.
Published: (2024)
by: Kim, Seok-Jin, et al.
Published: (2024)
Regret Bounds for Adversarial Contextual Bandits with General Function Approximation and Delayed Feedback
by: Levy, Orin, et al.
Published: (2025)
by: Levy, Orin, et al.
Published: (2025)
Online Estimation via Offline Estimation: An Information-Theoretic Framework
by: Foster, Dylan J., et al.
Published: (2024)
by: Foster, Dylan J., et al.
Published: (2024)
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
by: Liu, Hongyi, et al.
Published: (2025)
by: Liu, Hongyi, et al.
Published: (2025)
Near-Optimal Regret in Adversarial Kernel Bandits
by: Zhang, Yu-Jie, et al.
Published: (2026)
by: Zhang, Yu-Jie, et al.
Published: (2026)
Regret Distribution in Stochastic Bandits: Optimal Trade-off between Expectation and Tail Risk
by: Simchi-Levi, David, et al.
Published: (2023)
by: Simchi-Levi, David, et al.
Published: (2023)
Improved Regret Analysis in Gaussian Process Bandits: Optimality for Noiseless Reward, RKHS norm, and Non-Stationary Variance
by: Iwazaki, Shogo, et al.
Published: (2025)
by: Iwazaki, Shogo, et al.
Published: (2025)
Decentralized Contextual Bandits with Network Adaptivity
by: Deng, Chuyun, et al.
Published: (2025)
by: Deng, Chuyun, et al.
Published: (2025)
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
An Improved Algorithm for Adversarial Linear Contextual Bandits via Reduction
by: van Erven, Tim, et al.
Published: (2025)
by: van Erven, Tim, et al.
Published: (2025)
Bayesian Regret Minimization in Offline Bandits
by: Petrik, Marek, et al.
Published: (2023)
by: Petrik, Marek, et al.
Published: (2023)
Optimal Regret for Single Index Bandits
by: Dey, Devdan, et al.
Published: (2026)
by: Dey, Devdan, et al.
Published: (2026)
No-Regret Linear Bandits under Gap-Adjusted Misspecification
by: Liu, Chong, et al.
Published: (2025)
by: Liu, Chong, et al.
Published: (2025)
No-Regret is not enough! Bandits with General Constraints through Adaptive Regret Minimization
by: Bernasconi, Martino, et al.
Published: (2024)
by: Bernasconi, Martino, et al.
Published: (2024)
Cramming Contextual Bandits for On-policy Statistical Evaluation
by: Jia, Zeyang, et al.
Published: (2024)
by: Jia, Zeyang, et al.
Published: (2024)
Calibrated Recommendations with Contextual Bandits
by: Feijer, Diego, et al.
Published: (2025)
by: Feijer, Diego, et al.
Published: (2025)
Multi-User Contextual Cascading Bandits for Personalized Recommendation
by: Park, Jiho, et al.
Published: (2025)
by: Park, Jiho, et al.
Published: (2025)
Similar Items
-
On the Minimax Regret of Sequential Probability Assignment via Square-Root Entropy
by: Jia, Zeyu, et al.
Published: (2025) -
Self-Normalized Martingales and Uniform Regret Bounds for Linear Regression
by: Chen, Fan, et al.
Published: (2026) -
Variance-Dependent Regret Lower Bounds for Contextual Bandits
by: He, Jiafan, et al.
Published: (2025) -
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
by: Di, Qiwei, et al.
Published: (2023) -
Achieving Optimal Static and Dynamic Regret Simultaneously in Bandits with Deterministic Losses
by: Qian, Jian, et al.
Published: (2026)