Deceptive Exploration in Multi-armed Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Vurankaya, I. Arda, Karabag, Mustafa O., Suttle, Wesley A., Milzman, Jesse, Fridovich-Keil, David, Topcu, Ufuk |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Value of Information-based Deceptive Path Planning Under Adversarial Interventions
by: Suttle, Wesley A., et al.
Published: (2025)
by: Suttle, Wesley A., et al.
Published: (2025)
A Multi-Fidelity Control Variate Approach for Policy Gradient Estimation
by: Liu, Xinjie, et al.
Published: (2025)
by: Liu, Xinjie, et al.
Published: (2025)
Deceptive Planning Exploiting Inattention Blindness
by: Karabag, Mustafa O., et al.
Published: (2025)
by: Karabag, Mustafa O., et al.
Published: (2025)
Monotonic Transformation Invariant Multi-task Learning
by: Murthy, Surya, et al.
Published: (2025)
by: Murthy, Surya, et al.
Published: (2025)
Do LLMs Strategically Reveal, Conceal, and Infer Information? A Theoretical and Empirical Analysis in The Chameleon Game
by: Karabag, Mustafa O., et al.
Published: (2025)
by: Karabag, Mustafa O., et al.
Published: (2025)
When Should a Leader Act Suboptimally? The Role of Inferability in Repeated Stackelberg Games
by: Karabag, Mustafa O., et al.
Published: (2023)
by: Karabag, Mustafa O., et al.
Published: (2023)
Generalized Information Gathering Under Dynamics Uncertainty
by: Palafox, Fernando, et al.
Published: (2026)
by: Palafox, Fernando, et al.
Published: (2026)
Sequential Resource Trading Using Comparison-Based Gradient Estimation
by: Murthy, Surya, et al.
Published: (2024)
by: Murthy, Surya, et al.
Published: (2024)
Approximate Feedback Nash Equilibria with Sparse Inter-Agent Dependencies
by: Liu, Xinjie, et al.
Published: (2024)
by: Liu, Xinjie, et al.
Published: (2024)
Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions
by: Sobotka, Jan, et al.
Published: (2026)
by: Sobotka, Jan, et al.
Published: (2026)
Causally Abstracted Multi-armed Bandits
by: Zennaro, Fabio Massimo, et al.
Published: (2024)
by: Zennaro, Fabio Massimo, et al.
Published: (2024)
Linear-Quadratic Gaussian Games with Distributed Sparse Estimation
by: Qiu, Tianyu, et al.
Published: (2026)
by: Qiu, Tianyu, et al.
Published: (2026)
A Decentralized Shotgun Approach for Team Deception
by: Probine, Caleb, et al.
Published: (2024)
by: Probine, Caleb, et al.
Published: (2024)
Zero-Shot Reinforcement Learning via Function Encoders
by: Ingebrand, Tyler, et al.
Published: (2024)
by: Ingebrand, Tyler, et al.
Published: (2024)
Bayesian Inverse Games with High-Dimensional Multi-Modal Observations
by: Jain, Yash, et al.
Published: (2026)
by: Jain, Yash, et al.
Published: (2026)
Cooperative Bargaining Games Without Utilities: Mediated Solutions from Direction Oracles
by: Gupta, Kushagra, et al.
Published: (2025)
by: Gupta, Kushagra, et al.
Published: (2025)
Deceptive Sequential Decision-Making via Regularized Policy Optimization
by: Kim, Yerin, et al.
Published: (2025)
by: Kim, Yerin, et al.
Published: (2025)
Neural Port-Hamiltonian Differential Algebraic Equations for Compositional Learning of Electrical Networks
by: Neary, Cyrus, et al.
Published: (2024)
by: Neary, Cyrus, et al.
Published: (2024)
Adaptive Shielding for Safe Reinforcement Learning under Hidden-Parameter Dynamics Shifts
by: Kwon, Minjae, et al.
Published: (2025)
by: Kwon, Minjae, et al.
Published: (2025)
Towards a Pretrained Model for Restless Bandits via Multi-arm Generalization
by: Zhao, Yunfan, et al.
Published: (2023)
by: Zhao, Yunfan, et al.
Published: (2023)
Improving Thompson Sampling via Information Relaxation for Budgeted Multi-armed Bandits
by: Jeong, Woojin, et al.
Published: (2024)
by: Jeong, Woojin, et al.
Published: (2024)
TIGER-MARL: Enhancing Multi-Agent Reinforcement Learning with Temporal Information through Graph-based Embeddings and Representations
by: Gupta, Nikunj, et al.
Published: (2025)
by: Gupta, Nikunj, et al.
Published: (2025)
Sensing Resource Allocation Against Data-Poisoning Attacks in Traffic Routing
by: Yu, Yue, et al.
Published: (2024)
by: Yu, Yue, et al.
Published: (2024)
Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing
by: Zongo, Alex, et al.
Published: (2026)
by: Zongo, Alex, et al.
Published: (2026)
Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective
by: Hu, Xiao, et al.
Published: (2026)
by: Hu, Xiao, et al.
Published: (2026)
Confidence-Controlled Exploration: Efficient Sparse-Reward Policy Learning for Robot Navigation
by: Patel, Bhrij, et al.
Published: (2023)
by: Patel, Bhrij, et al.
Published: (2023)
Combinatorial Multi-armed Bandits: Arm Selection via Group Testing
by: Mukherjee, Arpan, et al.
Published: (2024)
by: Mukherjee, Arpan, et al.
Published: (2024)
Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments
by: Zhang, Ziyuan, et al.
Published: (2025)
by: Zhang, Ziyuan, et al.
Published: (2025)
Meritocratic Fairness in Budgeted Combinatorial Multi-armed Bandits via Shapley Values
by: Sharma, Shradha, et al.
Published: (2026)
by: Sharma, Shradha, et al.
Published: (2026)
A Flow Matching Algorithm for Many-Shot Adaptation to Unseen Distributions
by: Ingebrand, Tyler, et al.
Published: (2026)
by: Ingebrand, Tyler, et al.
Published: (2026)
Co-Exploration and Co-Exploitation via Shared Structure in Multi-Task Bandits
by: Mukherjee, Sumantrak, et al.
Published: (2025)
by: Mukherjee, Sumantrak, et al.
Published: (2025)
Incentivized Exploration of Non-Stationary Stochastic Bandits
by: Chakraborty, Sourav, et al.
Published: (2024)
by: Chakraborty, Sourav, et al.
Published: (2024)
On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits
by: Hou, Yunlong, et al.
Published: (2026)
by: Hou, Yunlong, et al.
Published: (2026)
Online Foundation Model Selection in Robotics
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
Exploitation Is All You Need... for Exploration
by: Rentschler, Micah, et al.
Published: (2025)
by: Rentschler, Micah, et al.
Published: (2025)
Auto-Encoding Bayesian Inverse Games
by: Liu, Xinjie, et al.
Published: (2024)
by: Liu, Xinjie, et al.
Published: (2024)
Active Inverse Learning in Stackelberg Trajectory Games
by: Ward, William, et al.
Published: (2023)
by: Ward, William, et al.
Published: (2023)
Using Large Language Models to Automate and Expedite Reinforcement Learning with Reward Machine
by: Alsadat, Shayan Meshkat, et al.
Published: (2024)
by: Alsadat, Shayan Meshkat, et al.
Published: (2024)
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
by: Kumarappan, Adarsh, et al.
Published: (2025)
by: Kumarappan, Adarsh, et al.
Published: (2025)
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
Similar Items
-
Value of Information-based Deceptive Path Planning Under Adversarial Interventions
by: Suttle, Wesley A., et al.
Published: (2025) -
A Multi-Fidelity Control Variate Approach for Policy Gradient Estimation
by: Liu, Xinjie, et al.
Published: (2025) -
Deceptive Planning Exploiting Inattention Blindness
by: Karabag, Mustafa O., et al.
Published: (2025) -
Monotonic Transformation Invariant Multi-task Learning
by: Murthy, Surya, et al.
Published: (2025) -
Do LLMs Strategically Reveal, Conceal, and Infer Information? A Theoretical and Empirical Analysis in The Chameleon Game
by: Karabag, Mustafa O., et al.
Published: (2025)