Multi-agent Multi-armed Bandits with Stochastic Sharable Arm Capacities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Hong, Mo, Jinyu, Lian, Defu, Wang, Jie, Chen, Enhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multiple-play Stochastic Bandits with Prioritized Arm Capacity Sharing
von: Xie, Hong, et al.
Veröffentlicht: (2025)
von: Xie, Hong, et al.
Veröffentlicht: (2025)
Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective
von: Hu, Xiao, et al.
Veröffentlicht: (2026)
von: Hu, Xiao, et al.
Veröffentlicht: (2026)
Federated Contextual Cascading Bandits with Asynchronous Communication and Heterogeneous Users
von: Yang, Hantao, et al.
Veröffentlicht: (2024)
von: Yang, Hantao, et al.
Veröffentlicht: (2024)
Analytical and Empirical Study of Herding Effects in Recommendation Systems
von: Xie, Hong, et al.
Veröffentlicht: (2024)
von: Xie, Hong, et al.
Veröffentlicht: (2024)
Combinatorial Multi-armed Bandits: Arm Selection via Group Testing
von: Mukherjee, Arpan, et al.
Veröffentlicht: (2024)
von: Mukherjee, Arpan, et al.
Veröffentlicht: (2024)
Demystifying Design Choices of Reinforcement Fine-tuning: A Batched Contextual Bandit Learning Perspective
von: Xie, Hong, et al.
Veröffentlicht: (2026)
von: Xie, Hong, et al.
Veröffentlicht: (2026)
Securing Recommender System via Cooperative Training
von: Wang, Qingyang, et al.
Veröffentlicht: (2024)
von: Wang, Qingyang, et al.
Veröffentlicht: (2024)
Model Specific Task Similarity for Vision Language Model Selection via Layer Conductance
von: Yang, Wei, et al.
Veröffentlicht: (2026)
von: Yang, Wei, et al.
Veröffentlicht: (2026)
Causally Abstracted Multi-armed Bandits
von: Zennaro, Fabio Massimo, et al.
Veröffentlicht: (2024)
von: Zennaro, Fabio Massimo, et al.
Veröffentlicht: (2024)
Deceptive Exploration in Multi-armed Bandits
von: Vurankaya, I. Arda, et al.
Veröffentlicht: (2025)
von: Vurankaya, I. Arda, et al.
Veröffentlicht: (2025)
Efficient Machine Unlearning via Influence Approximation
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
Networked Restless Multi-Arm Bandits with Reinforcement Learning
von: Zhang, Hanmo, et al.
Veröffentlicht: (2025)
von: Zhang, Hanmo, et al.
Veröffentlicht: (2025)
Understanding the planning of LLM agents: A survey
von: Huang, Xu, et al.
Veröffentlicht: (2024)
von: Huang, Xu, et al.
Veröffentlicht: (2024)
Denoising Pre-Training and Customized Prompt Learning for Efficient Multi-Behavior Sequential Recommendation
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Competitive Multi-armed Bandit Games for Resource Sharing
von: Li, Hongbo, et al.
Veröffentlicht: (2025)
von: Li, Hongbo, et al.
Veröffentlicht: (2025)
Foundations and Frontiers of Graph Learning Theory
von: Huang, Yu, et al.
Veröffentlicht: (2024)
von: Huang, Yu, et al.
Veröffentlicht: (2024)
Task Selection and Assignment for Multi-modal Multi-task Dialogue Act Classification with Non-stationary Multi-armed Bandits
von: He, Xiangheng, et al.
Veröffentlicht: (2023)
von: He, Xiangheng, et al.
Veröffentlicht: (2023)
PiXTime: A Model for Federated Time Series Forecasting with Heterogeneous Data across Nodes
von: Zhou, Yiming, et al.
Veröffentlicht: (2026)
von: Zhou, Yiming, et al.
Veröffentlicht: (2026)
Learning Complete Topology-Aware Correlations Between Relations for Inductive Link Prediction
von: Wang, Jie, et al.
Veröffentlicht: (2023)
von: Wang, Jie, et al.
Veröffentlicht: (2023)
LLM Cache Bandit Revisited: Addressing Query Heterogeneity for Cost-Effective LLM Inference
von: Yang, Hantao, et al.
Veröffentlicht: (2025)
von: Yang, Hantao, et al.
Veröffentlicht: (2025)
Improving Thompson Sampling via Information Relaxation for Budgeted Multi-armed Bandits
von: Jeong, Woojin, et al.
Veröffentlicht: (2024)
von: Jeong, Woojin, et al.
Veröffentlicht: (2024)
Towards a Pretrained Model for Restless Bandits via Multi-arm Generalization
von: Zhao, Yunfan, et al.
Veröffentlicht: (2023)
von: Zhao, Yunfan, et al.
Veröffentlicht: (2023)
Multi-armed Bandit and Backbone boost Lin-Kernighan-Helsgaun Algorithm for the Traveling Salesman Problems
von: Wang, Long, et al.
Veröffentlicht: (2025)
von: Wang, Long, et al.
Veröffentlicht: (2025)
Meritocratic Fairness in Budgeted Combinatorial Multi-armed Bandits via Shapley Values
von: Sharma, Shradha, et al.
Veröffentlicht: (2026)
von: Sharma, Shradha, et al.
Veröffentlicht: (2026)
HyperTree Planning: Enhancing LLM Reasoning via Hierarchical Thinking
von: Gui, Runquan, et al.
Veröffentlicht: (2025)
von: Gui, Runquan, et al.
Veröffentlicht: (2025)
WESE: Weak Exploration to Strong Exploitation for LLM Agents
von: Huang, Xu, et al.
Veröffentlicht: (2024)
von: Huang, Xu, et al.
Veröffentlicht: (2024)
M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
von: Chen, Jianlv, et al.
Veröffentlicht: (2024)
von: Chen, Jianlv, et al.
Veröffentlicht: (2024)
Best-Arm Identification in Unimodal Bandits
von: Poiani, Riccardo, et al.
Veröffentlicht: (2024)
von: Poiani, Riccardo, et al.
Veröffentlicht: (2024)
Contextual Restless Multi-Armed Bandits with Application to Demand Response Decision-Making
von: Chen, Xin, et al.
Veröffentlicht: (2024)
von: Chen, Xin, et al.
Veröffentlicht: (2024)
UniMEL: A Unified Framework for Multimodal Entity Linking with Large Language Models
von: Qi, Liu, et al.
Veröffentlicht: (2024)
von: Qi, Liu, et al.
Veröffentlicht: (2024)
Understanding Privacy Risks of Embeddings Induced by Large Language Models
von: Zhu, Zhihao, et al.
Veröffentlicht: (2024)
von: Zhu, Zhihao, et al.
Veröffentlicht: (2024)
Learning from Emptiness: De-biasing Listwise Rerankers with Content-Agnostic Probability Calibration
von: Lv, Hang, et al.
Veröffentlicht: (2026)
von: Lv, Hang, et al.
Veröffentlicht: (2026)
A Unified Frequency Domain Decomposition Framework for Interpretable and Robust Time Series Forecasting
von: He, Cheng, et al.
Veröffentlicht: (2025)
von: He, Cheng, et al.
Veröffentlicht: (2025)
Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments
von: Zhang, Ziyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyuan, et al.
Veröffentlicht: (2025)
Heterogeneous Multi-agent Multi-armed Bandits on Stochastic Block Models
von: Xu, Mengfan, et al.
Veröffentlicht: (2025)
von: Xu, Mengfan, et al.
Veröffentlicht: (2025)
Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback
von: Zeng, Qirun, et al.
Veröffentlicht: (2026)
von: Zeng, Qirun, et al.
Veröffentlicht: (2026)
Flickering Multi-Armed Bandits
von: Chakraborty, Sourav, et al.
Veröffentlicht: (2026)
von: Chakraborty, Sourav, et al.
Veröffentlicht: (2026)
Learning Partially Aligned Item Representation for Cross-Domain Sequential Recommendation
von: Yin, Mingjia, et al.
Veröffentlicht: (2024)
von: Yin, Mingjia, et al.
Veröffentlicht: (2024)
Influential Bandits: Pulling an Arm May Change the Environment
von: Sato, Ryoma, et al.
Veröffentlicht: (2025)
von: Sato, Ryoma, et al.
Veröffentlicht: (2025)
Rethinking Purity and Diversity in Multi-Behavior Sequential Recommendation from the Frequency Perspective
von: Han, Yongqiang, et al.
Veröffentlicht: (2025)
von: Han, Yongqiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multiple-play Stochastic Bandits with Prioritized Arm Capacity Sharing
von: Xie, Hong, et al.
Veröffentlicht: (2025) -
Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective
von: Hu, Xiao, et al.
Veröffentlicht: (2026) -
Federated Contextual Cascading Bandits with Asynchronous Communication and Heterogeneous Users
von: Yang, Hantao, et al.
Veröffentlicht: (2024) -
Analytical and Empirical Study of Herding Effects in Recommendation Systems
von: Xie, Hong, et al.
Veröffentlicht: (2024) -
Combinatorial Multi-armed Bandits: Arm Selection via Group Testing
von: Mukherjee, Arpan, et al.
Veröffentlicht: (2024)