Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Chenlu, Jin, Yujia, Agarwal, Alekh, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Corruption-Robust Algorithms with Uncertainty Weighting for Nonlinear Contextual Bandits and Markov Decision Processes
von: Ye, Chenlu, et al.
Veröffentlicht: (2022)
von: Ye, Chenlu, et al.
Veröffentlicht: (2022)
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
von: Zhao, Heyang, et al.
Veröffentlicht: (2024)
von: Zhao, Heyang, et al.
Veröffentlicht: (2024)
Low-rank Matrix Bandits with Heavy-tailed Rewards
von: Kang, Yue, et al.
Veröffentlicht: (2024)
von: Kang, Yue, et al.
Veröffentlicht: (2024)
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
von: Agarwal, Alekh, et al.
Veröffentlicht: (2022)
von: Agarwal, Alekh, et al.
Veröffentlicht: (2022)
Catoni-Style Change Point Detection for Regret Minimization in Non-Stationary Heavy-Tailed Bandits
von: Genalti, Gianmarco, et al.
Veröffentlicht: (2025)
von: Genalti, Gianmarco, et al.
Veröffentlicht: (2025)
Stochastic Gradient Succeeds for Bandits
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
Heavy-tailed Linear Bandits: Adversarial Robustness, Best-of-both-worlds, and Beyond
von: Zhao, Canzhe, et al.
Veröffentlicht: (2025)
von: Zhao, Canzhe, et al.
Veröffentlicht: (2025)
Single Index Bandits: Generalized Linear Contextual Bandits with Unknown Reward Functions
von: Kang, Yue, et al.
Veröffentlicht: (2025)
von: Kang, Yue, et al.
Veröffentlicht: (2025)
Robust Preference Optimization through Reward Model Distillation
von: Fisch, Adam, et al.
Veröffentlicht: (2024)
von: Fisch, Adam, et al.
Veröffentlicht: (2024)
Corruption-Robust Offline Reinforcement Learning with General Function Approximation
von: Ye, Chenlu, et al.
Veröffentlicht: (2023)
von: Ye, Chenlu, et al.
Veröffentlicht: (2023)
Towards Robust Model-Based Reinforcement Learning Against Adversarial Corruption
von: Ye, Chenlu, et al.
Veröffentlicht: (2024)
von: Ye, Chenlu, et al.
Veröffentlicht: (2024)
Group-Sensitive Offline Contextual Bandits
von: Guo, Yihong, et al.
Veröffentlicht: (2025)
von: Guo, Yihong, et al.
Veröffentlicht: (2025)
Robust and Computationally Efficient Linear Contextual Bandits under Adversarial Corruption and Heavy-Tailed Noise
von: Tani, Naoto, et al.
Veröffentlicht: (2026)
von: Tani, Naoto, et al.
Veröffentlicht: (2026)
Multi-agent Multi-armed Bandit with Fully Heavy-tailed Dynamics
von: Wang, Xingyu, et al.
Veröffentlicht: (2025)
von: Wang, Xingyu, et al.
Veröffentlicht: (2025)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
Signature Approach for Contextual Bandits with Nonlinear and Path-dependent Rewards
von: Guo, Xin, et al.
Veröffentlicht: (2026)
von: Guo, Xin, et al.
Veröffentlicht: (2026)
Improved Regret Bounds for Linear Bandits with Heavy-Tailed Rewards
von: Tajdini, Artin, et al.
Veröffentlicht: (2025)
von: Tajdini, Artin, et al.
Veröffentlicht: (2025)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
From Contextual Combinatorial Semi-Bandits to Bandit List Classification: Improved Sample Complexity with Sparse Rewards
von: Erez, Liad, et al.
Veröffentlicht: (2025)
von: Erez, Liad, et al.
Veröffentlicht: (2025)
Robust Offline Reinforcement learning with Heavy-Tailed Rewards
von: Zhu, Jin, et al.
Veröffentlicht: (2023)
von: Zhu, Jin, et al.
Veröffentlicht: (2023)
Wasserstein Distributionally Robust Policy Evaluation and Learning for Contextual Bandits
von: Shen, Yi, et al.
Veröffentlicht: (2023)
von: Shen, Yi, et al.
Veröffentlicht: (2023)
Design Considerations in Offline Preference-based RL
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
von: Agarwal, Alekh, et al.
Veröffentlicht: (2025)
Offline Imitation Learning from Multiple Baselines with Applications to Compiler Optimization
von: Marinov, Teodor V., et al.
Veröffentlicht: (2024)
von: Marinov, Teodor V., et al.
Veröffentlicht: (2024)
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
Inverse Contextual Bandits without Rewards: Learning from a Non-Stationary Learner via Suffix Imitation
von: Kong, Yuqi, et al.
Veröffentlicht: (2026)
von: Kong, Yuqi, et al.
Veröffentlicht: (2026)
Contextual Linear Bandits with Delay as Payoff
von: Zhang, Mengxiao, et al.
Veröffentlicht: (2025)
von: Zhang, Mengxiao, et al.
Veröffentlicht: (2025)
Contextual Bandits with Non-Stationary Correlated Rewards for User Association in MmWave Vehicular Networks
von: He, Xiaoyang, et al.
Veröffentlicht: (2024)
von: He, Xiaoyang, et al.
Veröffentlicht: (2024)
Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation
von: Chen, Ziru, et al.
Veröffentlicht: (2026)
von: Chen, Ziru, et al.
Veröffentlicht: (2026)
RIE-Greedy: Regularization-Induced Exploration for Contextual Bandits
von: Li, Tong, et al.
Veröffentlicht: (2026)
von: Li, Tong, et al.
Veröffentlicht: (2026)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2023)
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2023)
Efficient Contextual Bandits with Uninformed Feedback Graphs
von: Zhang, Mengxiao, et al.
Veröffentlicht: (2024)
von: Zhang, Mengxiao, et al.
Veröffentlicht: (2024)
Contextual Bandits for Unbounded Context Distributions
von: Zhao, Puning, et al.
Veröffentlicht: (2024)
von: Zhao, Puning, et al.
Veröffentlicht: (2024)
Sparse Nonparametric Contextual Bandits
von: Flynn, Hamish, et al.
Veröffentlicht: (2025)
von: Flynn, Hamish, et al.
Veröffentlicht: (2025)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
Mitigating Preference Hacking in Policy Optimization with Pessimism
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
von: Gupta, Dhawal, et al.
Veröffentlicht: (2025)
Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
Bandit Simulation for Average Reward Inference
von: Praharaj, Samya, et al.
Veröffentlicht: (2026)
von: Praharaj, Samya, et al.
Veröffentlicht: (2026)
BanditQ: Fair Bandits with Guaranteed Rewards
von: Sinha, Abhishek
Veröffentlicht: (2023)
von: Sinha, Abhishek
Veröffentlicht: (2023)
Distributionally Robust Policy Evaluation under General Covariate Shift in Contextual Bandits
von: Guo, Yihong, et al.
Veröffentlicht: (2024)
von: Guo, Yihong, et al.
Veröffentlicht: (2024)
Efficient Algorithms for Logistic Contextual Slate Bandits with Bandit Feedback
von: Goyal, Tanmay, et al.
Veröffentlicht: (2025)
von: Goyal, Tanmay, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Corruption-Robust Algorithms with Uncertainty Weighting for Nonlinear Contextual Bandits and Markov Decision Processes
von: Ye, Chenlu, et al.
Veröffentlicht: (2022) -
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
von: Zhao, Heyang, et al.
Veröffentlicht: (2024) -
Low-rank Matrix Bandits with Heavy-tailed Rewards
von: Kang, Yue, et al.
Veröffentlicht: (2024) -
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
von: Agarwal, Alekh, et al.
Veröffentlicht: (2022) -
Catoni-Style Change Point Detection for Regret Minimization in Non-Stationary Heavy-Tailed Bandits
von: Genalti, Gianmarco, et al.
Veröffentlicht: (2025)