Annealed Softmax Greedy in Many-Armed Bayesian Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Overman, William, Bayati, Mohsen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
by: Overman, William, et al.
Published: (2025)
by: Overman, William, et al.
Published: (2025)
The Unreasonable Effectiveness of Greedy Algorithms in Multi-Armed Bandit with Many Arms
by: Bayati, Mohsen, et al.
Published: (2020)
by: Bayati, Mohsen, et al.
Published: (2020)
Calibrating Conservatism for Scalable Oversight
by: Overman, William, et al.
Published: (2026)
by: Overman, William, et al.
Published: (2026)
Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language Models
by: Overman, William, et al.
Published: (2025)
by: Overman, William, et al.
Published: (2025)
Causal Effects with Unobserved Unit Types in Interacting Human-AI Systems
by: Overman, William, et al.
Published: (2026)
by: Overman, William, et al.
Published: (2026)
Flickering Multi-Armed Bandits
by: Chakraborty, Sourav, et al.
Published: (2026)
by: Chakraborty, Sourav, et al.
Published: (2026)
Aligning Model Properties via Conformal Risk Control
by: Overman, William, et al.
Published: (2024)
by: Overman, William, et al.
Published: (2024)
Multi-Armed Bandits With Best-Action Queries
by: Bacchiocchi, Francesco, et al.
Published: (2026)
by: Bacchiocchi, Francesco, et al.
Published: (2026)
Large Language Model-Enhanced Multi-Armed Bandits
by: Sun, Jiahang, et al.
Published: (2025)
by: Sun, Jiahang, et al.
Published: (2025)
Multi-Armed Bandits-Based Optimization of Decision Trees
by: Shanto, Hasibul Karim, et al.
Published: (2025)
by: Shanto, Hasibul Karim, et al.
Published: (2025)
Speed Up the Cold-Start Learning in Two-Sided Bandits with Many Arms
by: Bayati, Mohsen, et al.
Published: (2022)
by: Bayati, Mohsen, et al.
Published: (2022)
Introduction to Multi-Armed Bandits
by: Slivkins, Aleksandrs
Published: (2019)
by: Slivkins, Aleksandrs
Published: (2019)
When Exploration Comes for Free with Mixture-Greedy: Do we need UCB in Diversity-Aware Multi-Armed Bandits?
by: Nia, Bahar Dibaei, et al.
Published: (2026)
by: Nia, Bahar Dibaei, et al.
Published: (2026)
Fair Algorithms with Probing for Multi-Agent Multi-Armed Bandits
by: Xu, Tianyi, et al.
Published: (2025)
by: Xu, Tianyi, et al.
Published: (2025)
Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
Global Rewards in Restless Multi-Armed Bandits
by: Raman, Naveen, et al.
Published: (2024)
by: Raman, Naveen, et al.
Published: (2024)
IRL for Restless Multi-Armed Bandits with Applications in Maternal and Child Health
by: Jain, Gauri, et al.
Published: (2024)
by: Jain, Gauri, et al.
Published: (2024)
Thresholding Data Shapley for Data Cleansing Using Multi-Armed Bandits
by: Namba, Hiroyuki, et al.
Published: (2024)
by: Namba, Hiroyuki, et al.
Published: (2024)
Put CASH on Bandits: A Max K-Armed Problem for Automated Machine Learning
by: Balef, Amir Rezaei, et al.
Published: (2025)
by: Balef, Amir Rezaei, et al.
Published: (2025)
Mobility-Aware Federated Learning: Multi-Armed Bandit Based Selection in Vehicular Network
by: Tu, Haoyu, et al.
Published: (2024)
by: Tu, Haoyu, et al.
Published: (2024)
Online Prompt Pricing based on Combinatorial Multi-Armed Bandit and Hierarchical Stackelberg Game
by: Li, Meiling, et al.
Published: (2024)
by: Li, Meiling, et al.
Published: (2024)
ParBalans: Parallel Multi-Armed Bandits-based Adaptive Large Neighborhood Search
by: Yilmaz, Alican, et al.
Published: (2025)
by: Yilmaz, Alican, et al.
Published: (2025)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits
by: Hou, Yunlong, et al.
Published: (2026)
by: Hou, Yunlong, et al.
Published: (2026)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits
by: Ran, Dezhi, et al.
Published: (2025)
by: Ran, Dezhi, et al.
Published: (2025)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
by: Rezazadeh, Navid, et al.
Published: (2026)
by: Rezazadeh, Navid, et al.
Published: (2026)
Federated Combinatorial Multi-Agent Multi-Armed Bandits
by: Fourati, Fares, et al.
Published: (2024)
by: Fourati, Fares, et al.
Published: (2024)
Geometry-Aware Approaches for Balancing Performance and Theoretical Guarantees in Linear Bandits
by: Luo, Yuwei, et al.
Published: (2023)
by: Luo, Yuwei, et al.
Published: (2023)
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
by: Chen, Sanxing, et al.
Published: (2025)
by: Chen, Sanxing, et al.
Published: (2025)
Adaptive Budgeted Multi-Armed Bandits for IoT with Dynamic Resource Constraints
by: Vaishnav, Shubham, et al.
Published: (2025)
by: Vaishnav, Shubham, et al.
Published: (2025)
Higher-Order Causal Message Passing for Experimentation with Complex Interference
by: Bayati, Mohsen, et al.
Published: (2024)
by: Bayati, Mohsen, et al.
Published: (2024)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Bayesian Analysis of Combinatorial Gaussian Process Bandits
by: Sandberg, Jack, et al.
Published: (2023)
by: Sandberg, Jack, et al.
Published: (2023)
Can We Validate Counterfactual Estimations in the Presence of General Network Interference?
by: Shirani, Sadegh, et al.
Published: (2025)
by: Shirani, Sadegh, et al.
Published: (2025)
Decisions and Deployment: The Five-Year SAHELI Project (2020-2025) on Restless Multi-Armed Bandits for Improving Maternal and Child Health
by: Verma, Shresth, et al.
Published: (2026)
by: Verma, Shresth, et al.
Published: (2026)
Softmax is not Enough (for Adaptive Conformal Classification)
by: Attar, Navid Akhavan, et al.
Published: (2026)
by: Attar, Navid Akhavan, et al.
Published: (2026)
A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public Health
by: Behari, Nikhil, et al.
Published: (2024)
by: Behari, Nikhil, et al.
Published: (2024)
Balans: Multi-Armed Bandits-based Adaptive Large Neighborhood Search for Mixed-Integer Programming Problem
by: Cai, Junyang, et al.
Published: (2024)
by: Cai, Junyang, et al.
Published: (2024)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
Similar Items
-
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
by: Overman, William, et al.
Published: (2025) -
The Unreasonable Effectiveness of Greedy Algorithms in Multi-Armed Bandit with Many Arms
by: Bayati, Mohsen, et al.
Published: (2020) -
Calibrating Conservatism for Scalable Oversight
by: Overman, William, et al.
Published: (2026) -
Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language Models
by: Overman, William, et al.
Published: (2025) -
Causal Effects with Unobserved Unit Types in Interacting Human-AI Systems
by: Overman, William, et al.
Published: (2026)