$(ε, u)$-Adaptive Regret Minimization in Heavy-Tailed Bandits
Fuente:
arXiv
Guardado en:
| Autores principales: | Genalti, Gianmarco, Marsigli, Lupo, Gatti, Nicola, Metelli, Alberto Maria |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Catoni-Style Change Point Detection for Regret Minimization in Non-Stationary Heavy-Tailed Bandits
por: Genalti, Gianmarco, et al.
Publicado: (2025)
por: Genalti, Gianmarco, et al.
Publicado: (2025)
Bridging Rested and Restless Bandits with Graph-Triggering: Rising and Rotting
por: Genalti, Gianmarco, et al.
Publicado: (2024)
por: Genalti, Gianmarco, et al.
Publicado: (2024)
Autoregressive Bandits
por: Bacchiocchi, Francesco, et al.
Publicado: (2022)
por: Bacchiocchi, Francesco, et al.
Publicado: (2022)
Data-Dependent Regret Bounds for Constrained MABs
por: Genalti, Gianmarco, et al.
Publicado: (2025)
por: Genalti, Gianmarco, et al.
Publicado: (2025)
No-Regret Reinforcement Learning in Smooth MDPs
por: Maran, Davide, et al.
Publicado: (2024)
por: Maran, Davide, et al.
Publicado: (2024)
No-Regret Learning Under Adversarial Resource Constraints: A Spending Plan Is All You Need!
por: Stradi, Francesco Emanuele, et al.
Publicado: (2025)
por: Stradi, Francesco Emanuele, et al.
Publicado: (2025)
On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits
por: Hou, Yunlong, et al.
Publicado: (2026)
por: Hou, Yunlong, et al.
Publicado: (2026)
Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic Bandits
por: Xue, Bo, et al.
Publicado: (2025)
por: Xue, Bo, et al.
Publicado: (2025)
Adaptive Heavy-Tailed Stochastic Gradient Descent
por: Gong, Bodu, et al.
Publicado: (2025)
por: Gong, Bodu, et al.
Publicado: (2025)
Generalizing Behavior via Inverse Reinforcement Learning with Closed-Form Reward Centroids
por: Lazzati, Filippo, et al.
Publicado: (2025)
por: Lazzati, Filippo, et al.
Publicado: (2025)
Imitation Learning as Return Distribution Matching
por: Lazzati, Filippo, et al.
Publicado: (2025)
por: Lazzati, Filippo, et al.
Publicado: (2025)
Tackling Heavy-Tailed Rewards in Reinforcement Learning with Function Approximation: Minimax Optimal and Instance-Dependent Regret Bounds
por: Huang, Jiayi, et al.
Publicado: (2023)
por: Huang, Jiayi, et al.
Publicado: (2023)
Replicable Constrained Bandits
por: Bollini, Matteo, et al.
Publicado: (2026)
por: Bollini, Matteo, et al.
Publicado: (2026)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
por: He, Jiafan, et al.
Publicado: (2025)
por: He, Jiafan, et al.
Publicado: (2025)
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
por: Yadav, Robin, et al.
Publicado: (2025)
por: Yadav, Robin, et al.
Publicado: (2025)
Information Capacity Regret Bounds for Bandits with Mediator Feedback
por: Eldowa, Khaled, et al.
Publicado: (2024)
por: Eldowa, Khaled, et al.
Publicado: (2024)
Variance-Dependent Regret Bounds for Non-stationary Linear Bandits
por: Wang, Zhiyong, et al.
Publicado: (2024)
por: Wang, Zhiyong, et al.
Publicado: (2024)
A Best-of-Both-Worlds Algorithm for Constrained MDPs with Long-Term Constraints
por: Germano, Jacopo, et al.
Publicado: (2023)
por: Germano, Jacopo, et al.
Publicado: (2023)
Statistical Analysis of Policy Space Compression Problem
por: Molaei, Majid, et al.
Publicado: (2024)
por: Molaei, Majid, et al.
Publicado: (2024)
Projection by Convolution: Optimal Sample Complexity for Reinforcement Learning in Continuous-Space MDPs
por: Maran, Davide, et al.
Publicado: (2024)
por: Maran, Davide, et al.
Publicado: (2024)
Inverse Reinforcement Learning with Sub-optimal Experts
por: Poiani, Riccardo, et al.
Publicado: (2024)
por: Poiani, Riccardo, et al.
Publicado: (2024)
Generalized Kernelized Bandits: A Novel Self-Normalized Bernstein-Like Dimension-Free Inequality and Regret Bounds
por: Metelli, Alberto Maria, et al.
Publicado: (2025)
por: Metelli, Alberto Maria, et al.
Publicado: (2025)
TailedTS: Benchmark Dataset for Heavy-Tailed Time Series Prediction and Periodicity Quantification
por: Chen, Xinyu, et al.
Publicado: (2026)
por: Chen, Xinyu, et al.
Publicado: (2026)
Robust Offline Reinforcement learning with Heavy-Tailed Rewards
por: Zhu, Jin, et al.
Publicado: (2023)
por: Zhu, Jin, et al.
Publicado: (2023)
The Minimal Search Space for Conditional Causal Bandits
por: Simoes, Francisco N. F. Q., et al.
Publicado: (2025)
por: Simoes, Francisco N. F. Q., et al.
Publicado: (2025)
Improved Regret Bounds for Linear Bandits with Heavy-Tailed Rewards
por: Tajdini, Artin, et al.
Publicado: (2025)
por: Tajdini, Artin, et al.
Publicado: (2025)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
por: Ji, Kaixuan, et al.
Publicado: (2026)
por: Ji, Kaixuan, et al.
Publicado: (2026)
Adaptive Budget Optimization for Multichannel Advertising Using Combinatorial Bandits
por: Gangopadhyay, Briti, et al.
Publicado: (2025)
por: Gangopadhyay, Briti, et al.
Publicado: (2025)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
por: Hou, Yunlong, et al.
Publicado: (2025)
por: Hou, Yunlong, et al.
Publicado: (2025)
HTMuon: Improving Muon via Heavy-Tailed Spectral Correction
por: Pang, Tianyu, et al.
Publicado: (2026)
por: Pang, Tianyu, et al.
Publicado: (2026)
QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
por: Cho, Nicole, et al.
Publicado: (2025)
por: Cho, Nicole, et al.
Publicado: (2025)
Label-Efficient Monitoring of Classification Models via Stratified Importance Sampling
por: Marsigli, Lupo, et al.
Publicado: (2026)
por: Marsigli, Lupo, et al.
Publicado: (2026)
Satisficing Regret Minimization in Bandits: Constant Rate and Light-Tailed Distribution
por: Feng, Qing, et al.
Publicado: (2024)
por: Feng, Qing, et al.
Publicado: (2024)
Batch-Size Independent Regret Bounds for Combinatorial Semi-Bandits with Probabilistically Triggered Arms or Independent Arms
por: Liu, Xutong, et al.
Publicado: (2022)
por: Liu, Xutong, et al.
Publicado: (2022)
Causal Contextual Bandits with Adaptive Context
por: Madhavan, Rahul, et al.
Publicado: (2024)
por: Madhavan, Rahul, et al.
Publicado: (2024)
Markov Chain Decoders Overcome the Heavy-Tail Limitations of Lipschitz Generative Models
por: Ziani, Abdelhakim, et al.
Publicado: (2026)
por: Ziani, Abdelhakim, et al.
Publicado: (2026)
No-Regret is not enough! Bandits with General Constraints through Adaptive Regret Minimization
por: Bernasconi, Martino, et al.
Publicado: (2024)
por: Bernasconi, Martino, et al.
Publicado: (2024)
FOSSIL: Regret-Minimizing Curriculum Learning for Metadata-Free and Low-Data Mpox Diagnosis
por: Han, Sahng-Min, et al.
Publicado: (2025)
por: Han, Sahng-Min, et al.
Publicado: (2025)
Phase-Type Variational Autoencoders for Heavy-Tailed Data
por: Ziani, Abdelhakim, et al.
Publicado: (2026)
por: Ziani, Abdelhakim, et al.
Publicado: (2026)
Next-Token Prediction and Regret Minimization
por: Mohri, Mehryar, et al.
Publicado: (2026)
por: Mohri, Mehryar, et al.
Publicado: (2026)
Ejemplares similares
-
Catoni-Style Change Point Detection for Regret Minimization in Non-Stationary Heavy-Tailed Bandits
por: Genalti, Gianmarco, et al.
Publicado: (2025) -
Bridging Rested and Restless Bandits with Graph-Triggering: Rising and Rotting
por: Genalti, Gianmarco, et al.
Publicado: (2024) -
Autoregressive Bandits
por: Bacchiocchi, Francesco, et al.
Publicado: (2022) -
Data-Dependent Regret Bounds for Constrained MABs
por: Genalti, Gianmarco, et al.
Publicado: (2025) -
No-Regret Reinforcement Learning in Smooth MDPs
por: Maran, Davide, et al.
Publicado: (2024)