Saved in:
| Main Authors: | Ghosh, Ayon, Prashanth, L. A., Sen, Dipayan, Gopalan, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2203.16810 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Concentration Bounds for Optimized Certainty Equivalent Risk Estimation
by: Ghosh, Ayon, et al.
Published: (2024)
by: Ghosh, Ayon, et al.
Published: (2024)
Active clustering with bandit feedback
by: Thuot, Victor, et al.
Published: (2024)
by: Thuot, Victor, et al.
Published: (2024)
Trading off rewards and errors in multi-armed bandits
by: Erraqabi, Akram, et al.
Published: (2026)
by: Erraqabi, Akram, et al.
Published: (2026)
Fairness in two-player zero-sum games with bandit feedback
by: Akash, S, et al.
Published: (2026)
by: Akash, S, et al.
Published: (2026)
Towards Reliable Alignment: Uncertainty-aware RLHF
by: Banerjee, Debangshu, et al.
Published: (2024)
by: Banerjee, Debangshu, et al.
Published: (2024)
Bad Values but Good Behavior: Learning Highly Misspecified Bandits and MDPs
by: Banerjee, Debangshu, et al.
Published: (2023)
by: Banerjee, Debangshu, et al.
Published: (2023)
Optimal last-iterate convergence in matrix games with bandit feedback using the log-barrier
by: Fiegel, Come, et al.
Published: (2026)
by: Fiegel, Come, et al.
Published: (2026)
Extreme bandits
by: Carpentier, Alexandra, et al.
Published: (2026)
by: Carpentier, Alexandra, et al.
Published: (2026)
Testing the Feasibility of Linear Programs with Bandit Feedback
by: Gangrade, Aditya, et al.
Published: (2024)
by: Gangrade, Aditya, et al.
Published: (2024)
Stochastic contextual bandits with graph feedback: from independence number to MAS number
by: Wen, Yuxiao, et al.
Published: (2024)
by: Wen, Yuxiao, et al.
Published: (2024)
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Towards Reliable, Uncertainty-Aware Alignment
by: Banerjee, Debangshu, et al.
Published: (2025)
by: Banerjee, Debangshu, et al.
Published: (2025)
Why DPO is a Misspecified Estimator and How to Fix It
by: Gopalan, Aditya, et al.
Published: (2025)
by: Gopalan, Aditya, et al.
Published: (2025)
Approximate information maximization for bandit games
by: Barbier-Chebbah, Alex, et al.
Published: (2023)
by: Barbier-Chebbah, Alex, et al.
Published: (2023)
Does DQN Learn?
by: Gopalan, Aditya, et al.
Published: (2022)
by: Gopalan, Aditya, et al.
Published: (2022)
On the price of exact truthfulness in incentive-compatible online learning with bandit feedback: A regret lower bound for WSU-UX
by: Mortazavi, Ali, et al.
Published: (2024)
by: Mortazavi, Ali, et al.
Published: (2024)
Instance-dependent Stochastic Lipschitz bandit
by: Potfer, Marius, et al.
Published: (2026)
by: Potfer, Marius, et al.
Published: (2026)
Spectral bandits for smooth graph functions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Online learning in bandits with predicted context
by: Guo, Yongyi, et al.
Published: (2023)
by: Guo, Yongyi, et al.
Published: (2023)
Risk and optimal policies in bandit experiments
by: Adusumilli, Karun
Published: (2021)
by: Adusumilli, Karun
Published: (2021)
Revealing graph bandits for maximizing local influence
by: Carpentier, Alexandra, et al.
Published: (2026)
by: Carpentier, Alexandra, et al.
Published: (2026)
On the optimal regret of collaborative personalized linear bandits
by: Huang, Bruce, et al.
Published: (2025)
by: Huang, Bruce, et al.
Published: (2025)
Offline-to-online hyperparameter transfer for stochastic bandits
by: Sharma, Dravyansh, et al.
Published: (2025)
by: Sharma, Dravyansh, et al.
Published: (2025)
Linear bandits with polylogarithmic minimax regret
by: Lumbreras, Josep, et al.
Published: (2024)
by: Lumbreras, Josep, et al.
Published: (2024)
A policy gradient approach for optimization of smooth risk measures
by: Vijayan, Nithia, et al.
Published: (2022)
by: Vijayan, Nithia, et al.
Published: (2022)
Policy Gradient Methods for Distortion Risk Measures
by: Vijayan, Nithia, et al.
Published: (2021)
by: Vijayan, Nithia, et al.
Published: (2021)
Smoothed functional-based gradient algorithms for off-policy reinforcement learning: A non-asymptotic viewpoint
by: Vijayan, Nithia, et al.
Published: (2021)
by: Vijayan, Nithia, et al.
Published: (2021)
Lookahead identification in adversarial bandits: accuracy and memory bounds
by: Brukhim, Nataly, et al.
Published: (2026)
by: Brukhim, Nataly, et al.
Published: (2026)
Efficient kernelized bandit algorithms via exploration distributions
by: Hu, Bingshan, et al.
Published: (2025)
by: Hu, Bingshan, et al.
Published: (2025)
Leveraging priors on distribution functions for multi-arm bandits
by: Vashishtha, Sumit, et al.
Published: (2025)
by: Vashishtha, Sumit, et al.
Published: (2025)
Ensemble sampling for linear bandits: small ensembles suffice
by: Janz, David, et al.
Published: (2023)
by: Janz, David, et al.
Published: (2023)
VITS : Variational Inference Thompson Sampling for contextual bandits
by: Clavier, Pierre, et al.
Published: (2023)
by: Clavier, Pierre, et al.
Published: (2023)
When and why randomised exploration works (in linear bandits)
by: Abeille, Marc, et al.
Published: (2025)
by: Abeille, Marc, et al.
Published: (2025)
Automatic mixed precision for optimizing gained time with constrained loss mean-squared-error based on model partition to sequential sub-graphs
by: Markovich-Golan, Shmulik, et al.
Published: (2025)
by: Markovich-Golan, Shmulik, et al.
Published: (2025)
Adversarial bandit optimization for approximately linear functions
by: Cheng, Zhuoyu, et al.
Published: (2025)
by: Cheng, Zhuoyu, et al.
Published: (2025)
Adversarial Combinatorial Semi-bandits with Graph Feedback
by: Wen, Yuxiao
Published: (2025)
by: Wen, Yuxiao
Published: (2025)
Information-directed sampling for bandits: a primer
by: Hirling, Annika, et al.
Published: (2025)
by: Hirling, Annika, et al.
Published: (2025)
A survey on multi-player bandits
by: Boursier, Etienne, et al.
Published: (2022)
by: Boursier, Etienne, et al.
Published: (2022)
Information maximization for a broad variety of multi-armed bandit games
by: Barbier-Chebbah, Alex, et al.
Published: (2025)
by: Barbier-Chebbah, Alex, et al.
Published: (2025)
Optimizing Shortfall Risk Metric for Learning Regression Models
by: Ramaswamy, Harish G., et al.
Published: (2025)
by: Ramaswamy, Harish G., et al.
Published: (2025)
Similar Items
-
Concentration Bounds for Optimized Certainty Equivalent Risk Estimation
by: Ghosh, Ayon, et al.
Published: (2024) -
Active clustering with bandit feedback
by: Thuot, Victor, et al.
Published: (2024) -
Trading off rewards and errors in multi-armed bandits
by: Erraqabi, Akram, et al.
Published: (2026) -
Fairness in two-player zero-sum games with bandit feedback
by: Akash, S, et al.
Published: (2026) -
Towards Reliable Alignment: Uncertainty-aware RLHF
by: Banerjee, Debangshu, et al.
Published: (2024)