Pessimistic Risk-Aware Policy Learning in Contextual Bandits
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wan, Yilong, Li, Yuqiang, Wu, Xianyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PAC Off-Policy Prediction of Contextual Bandits
von: Wan, Yilong, et al.
Veröffentlicht: (2025)
von: Wan, Yilong, et al.
Veröffentlicht: (2025)
Improved Model-based Reinforcement Learning with Smooth Kernels
von: Long, Kun, et al.
Veröffentlicht: (2026)
von: Long, Kun, et al.
Veröffentlicht: (2026)
Risk-Aware Continuous Control with Neural Contextual Bandits
von: Ayala-Romero, Jose A., et al.
Veröffentlicht: (2023)
von: Ayala-Romero, Jose A., et al.
Veröffentlicht: (2023)
Learning a Pessimistic Reward Model in RLHF
von: Xu, Yinglun, et al.
Veröffentlicht: (2025)
von: Xu, Yinglun, et al.
Veröffentlicht: (2025)
Pessimistic Backward Policy for GFlowNets
von: Jang, Hyosoon, et al.
Veröffentlicht: (2024)
von: Jang, Hyosoon, et al.
Veröffentlicht: (2024)
Wasserstein Distributionally Robust Policy Evaluation and Learning for Contextual Bandits
von: Shen, Yi, et al.
Veröffentlicht: (2023)
von: Shen, Yi, et al.
Veröffentlicht: (2023)
Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and Learning
von: Sakhi, Otmane, et al.
Veröffentlicht: (2024)
von: Sakhi, Otmane, et al.
Veröffentlicht: (2024)
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits
von: Han, Zean, et al.
Veröffentlicht: (2026)
von: Han, Zean, et al.
Veröffentlicht: (2026)
Optimal Regret for Policy Optimization in Contextual Bandits
von: Levy, Orin, et al.
Veröffentlicht: (2026)
von: Levy, Orin, et al.
Veröffentlicht: (2026)
Pessimistic Off-Policy Optimization for Learning to Rank
von: Cief, Matej, et al.
Veröffentlicht: (2022)
von: Cief, Matej, et al.
Veröffentlicht: (2022)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024)
von: Zhai, Yuanzhao, et al.
Veröffentlicht: (2024)
Neural Risk-sensitive Satisficing in Contextual Bandits
von: Ito, Shogo, et al.
Veröffentlicht: (2025)
von: Ito, Shogo, et al.
Veröffentlicht: (2025)
Context-Action Embedding Learning for Off-Policy Evaluation in Contextual Bandits
von: Chandak, Kushagra, et al.
Veröffentlicht: (2025)
von: Chandak, Kushagra, et al.
Veröffentlicht: (2025)
Effective Off-Policy Evaluation and Learning in Contextual Combinatorial Bandits
von: Shimizu, Tatsuhiro, et al.
Veröffentlicht: (2024)
von: Shimizu, Tatsuhiro, et al.
Veröffentlicht: (2024)
Active Learning for Stochastic Contextual Linear Bandits
von: Brunskill, Emma, et al.
Veröffentlicht: (2026)
von: Brunskill, Emma, et al.
Veröffentlicht: (2026)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
von: Di, Qiwei, et al.
Veröffentlicht: (2023)
von: Di, Qiwei, et al.
Veröffentlicht: (2023)
Bayesian Inference of Contextual Bandit Policies via Empirical Likelihood
von: Ouyang, Jiangrong, et al.
Veröffentlicht: (2026)
von: Ouyang, Jiangrong, et al.
Veröffentlicht: (2026)
Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits
von: Li, Xuheng, et al.
Veröffentlicht: (2025)
von: Li, Xuheng, et al.
Veröffentlicht: (2025)
Optimistic Policy Learning under Pessimistic Adversaries with Regret and Violation Guarantees
von: Ganguly, Sourav, et al.
Veröffentlicht: (2026)
von: Ganguly, Sourav, et al.
Veröffentlicht: (2026)
POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes
von: Zhang, Ruijia, et al.
Veröffentlicht: (2025)
von: Zhang, Ruijia, et al.
Veröffentlicht: (2025)
Optimal Baseline Corrections for Off-Policy Contextual Bandits
von: Gupta, Shashank, et al.
Veröffentlicht: (2024)
von: Gupta, Shashank, et al.
Veröffentlicht: (2024)
Uniform Pessimistic Risk and its Optimal Portfolio
von: Hong, Sungchul, et al.
Veröffentlicht: (2023)
von: Hong, Sungchul, et al.
Veröffentlicht: (2023)
Transfer Learning for Contextual Multi-armed Bandits
von: Cai, Changxiao, et al.
Veröffentlicht: (2022)
von: Cai, Changxiao, et al.
Veröffentlicht: (2022)
Variance-Aware Linear UCB with Deep Representation for Neural Contextual Bandits
von: Bui, Ha Manh, et al.
Veröffentlicht: (2024)
von: Bui, Ha Manh, et al.
Veröffentlicht: (2024)
Latency-Aware Contextual Bandit: Application to Cryo-EM Data Collection
von: Wei, Lai, et al.
Veröffentlicht: (2024)
von: Wei, Lai, et al.
Veröffentlicht: (2024)
Learning When to Trust in Contextual Bandits
von: Ghasemi, Majid, et al.
Veröffentlicht: (2026)
von: Ghasemi, Majid, et al.
Veröffentlicht: (2026)
Optimizing Warfarin Dosing Using Contextual Bandit: An Offline Policy Learning and Evaluation Method
von: Huang, Yong, et al.
Veröffentlicht: (2024)
von: Huang, Yong, et al.
Veröffentlicht: (2024)
Distributionally Robust Policy Evaluation under General Covariate Shift in Contextual Bandits
von: Guo, Yihong, et al.
Veröffentlicht: (2024)
von: Guo, Yihong, et al.
Veröffentlicht: (2024)
Contextual Bandits for Unbounded Context Distributions
von: Zhao, Puning, et al.
Veröffentlicht: (2024)
von: Zhao, Puning, et al.
Veröffentlicht: (2024)
Sparse Nonparametric Contextual Bandits
von: Flynn, Hamish, et al.
Veröffentlicht: (2025)
von: Flynn, Hamish, et al.
Veröffentlicht: (2025)
PAK-UCB Contextual Bandit: An Online Learning Approach to Prompt-Aware Selection of Generative Models and LLMs
von: Hu, Xiaoyan, et al.
Veröffentlicht: (2024)
von: Hu, Xiaoyan, et al.
Veröffentlicht: (2024)
Feasibility-Aware Pessimistic Estimation: Toward Long-Horizon Safety in Offline RL
von: Tao, Zhikun
Veröffentlicht: (2025)
von: Tao, Zhikun
Veröffentlicht: (2025)
Efficient Algorithms for Logistic Contextual Slate Bandits with Bandit Feedback
von: Goyal, Tanmay, et al.
Veröffentlicht: (2025)
von: Goyal, Tanmay, et al.
Veröffentlicht: (2025)
Single Index Bandits: Generalized Linear Contextual Bandits with Unknown Reward Functions
von: Kang, Yue, et al.
Veröffentlicht: (2025)
von: Kang, Yue, et al.
Veröffentlicht: (2025)
Linear Contextual Bandits with Interference
von: Xu, Yang, et al.
Veröffentlicht: (2024)
von: Xu, Yang, et al.
Veröffentlicht: (2024)
Contextual Bandits for Resource-Constrained Devices using Probabilistic Learning
von: Angioli, Marco, et al.
Veröffentlicht: (2026)
von: Angioli, Marco, et al.
Veröffentlicht: (2026)
Learning with Incomplete Context: Linear Contextual Bandits with Pretrained Imputation
von: Yan, Hao, et al.
Veröffentlicht: (2025)
von: Yan, Hao, et al.
Veröffentlicht: (2025)
Episodic Contextual Bandits with Knapsacks under Conversion Models
von: Cheung, Wang Chi, et al.
Veröffentlicht: (2025)
von: Cheung, Wang Chi, et al.
Veröffentlicht: (2025)
Constrained Contextual Bandits with Adversarial Contexts
von: Sarkar, Dhruv, et al.
Veröffentlicht: (2026)
von: Sarkar, Dhruv, et al.
Veröffentlicht: (2026)
Neural Exploitation and Exploration of Contextual Bandits
von: Ban, Yikun, et al.
Veröffentlicht: (2023)
von: Ban, Yikun, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
PAC Off-Policy Prediction of Contextual Bandits
von: Wan, Yilong, et al.
Veröffentlicht: (2025) -
Improved Model-based Reinforcement Learning with Smooth Kernels
von: Long, Kun, et al.
Veröffentlicht: (2026) -
Risk-Aware Continuous Control with Neural Contextual Bandits
von: Ayala-Romero, Jose A., et al.
Veröffentlicht: (2023) -
Learning a Pessimistic Reward Model in RLHF
von: Xu, Yinglun, et al.
Veröffentlicht: (2025) -
Pessimistic Backward Policy for GFlowNets
von: Jang, Hyosoon, et al.
Veröffentlicht: (2024)