Offline Learning for Combinatorial Multi-armed Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xutong, Dai, Xiangxiang, Zuo, Jinhang, Wang, Siwei, Joe-Wong, Carlee, Lui, John C. S., Chen, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Batch-Size Independent Regret Bounds for Combinatorial Semi-Bandits with Probabilistically Triggered Arms or Independent Arms
by: Liu, Xutong, et al.
Published: (2022)
by: Liu, Xutong, et al.
Published: (2022)
Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
by: Liu, Xutong, et al.
Published: (2025)
by: Liu, Xutong, et al.
Published: (2025)
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
by: Liu, Xutong, et al.
Published: (2023)
by: Liu, Xutong, et al.
Published: (2023)
A Unified Online-Offline Framework for Co-Branding Campaign Recommendations
by: Dai, Xiangxiang, et al.
Published: (2025)
by: Dai, Xiangxiang, et al.
Published: (2025)
Combinatorial Logistic Bandits
by: Liu, Xutong, et al.
Published: (2024)
by: Liu, Xutong, et al.
Published: (2024)
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond
by: Liu, Xutong, et al.
Published: (2024)
by: Liu, Xutong, et al.
Published: (2024)
Continuous Semantic Caching for Low-Cost LLM Serving
by: Atalar, Baran, et al.
Published: (2026)
by: Atalar, Baran, et al.
Published: (2026)
Offline Clustering of Linear Bandits: The Power of Clusters under Limited Data
by: Liu, Jingyuan, et al.
Published: (2025)
by: Liu, Jingyuan, et al.
Published: (2025)
Stochastic Bandits Robust to Adversarial Attacks
by: Wang, Xuchuang, et al.
Published: (2024)
by: Wang, Xuchuang, et al.
Published: (2024)
Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution
by: Poon, Manhin, et al.
Published: (2025)
by: Poon, Manhin, et al.
Published: (2025)
Neural Combinatorial Clustered Bandits for Recommendation Systems
by: Atalar, Baran, et al.
Published: (2024)
by: Atalar, Baran, et al.
Published: (2024)
Offline Clustering of Preference Learning with Active-data Augmentation
by: Liu, Jingyuan, et al.
Published: (2025)
by: Liu, Jingyuan, et al.
Published: (2025)
Fusing Reward and Dueling Feedback in Stochastic Bandits
by: Wang, Xuchuang, et al.
Published: (2025)
by: Wang, Xuchuang, et al.
Published: (2025)
Intelligent Communication Planning for Constrained Environmental IoT Sensing with Reinforcement Learning
by: Hu, Yi, et al.
Published: (2023)
by: Hu, Yi, et al.
Published: (2023)
Tin-Tin: Towards Tiny Learning on Tiny Devices with Integer-based Neural Network Training
by: Hu, Yi, et al.
Published: (2025)
by: Hu, Yi, et al.
Published: (2025)
Cost-Ordered Feasibility for Multi-Armed Bandits with Cost Subsidy
by: Juneja, Ishank, et al.
Published: (2026)
by: Juneja, Ishank, et al.
Published: (2026)
Cost-Effective Online Multi-LLM Selection with Versatile Reward Models
by: Dai, Xiangxiang, et al.
Published: (2024)
by: Dai, Xiangxiang, et al.
Published: (2024)
Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE Inference
by: Han, Ziyi, et al.
Published: (2025)
by: Han, Ziyi, et al.
Published: (2025)
Learning with Limited Shared Information in Multi-agent Multi-armed Bandit
by: Shao, Junning, et al.
Published: (2025)
by: Shao, Junning, et al.
Published: (2025)
CoRAST: Towards Foundation Model-Powered Correlated Data Analysis in Resource-Constrained CPS and IoT
by: Hu, Yi, et al.
Published: (2024)
by: Hu, Yi, et al.
Published: (2024)
Heterogeneous Multi-agent Multi-armed Bandits on Stochastic Block Models
by: Xu, Mengfan, et al.
Published: (2025)
by: Xu, Mengfan, et al.
Published: (2025)
Hybrid Combinatorial Multi-armed Bandits with Probabilistically Triggered Arms
by: Zhou, Kongchang, et al.
Published: (2025)
by: Zhou, Kongchang, et al.
Published: (2025)
Combinatorial Rising Bandits
by: Song, Seockbean, et al.
Published: (2024)
by: Song, Seockbean, et al.
Published: (2024)
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
Leveraging the Power of Conversations: Optimal Key Term Selection in Conversational Contextual Bandits
by: Liu, Maoli, et al.
Published: (2025)
by: Liu, Maoli, et al.
Published: (2025)
Efficient and Optimal Policy Gradient Algorithm for Corrupted Multi-armed Bandits
by: Liu, Jiayuan, et al.
Published: (2025)
by: Liu, Jiayuan, et al.
Published: (2025)
Pairwise Elimination with Instance-Dependent Guarantees for Bandits with Cost Subsidy
by: Juneja, Ishank, et al.
Published: (2025)
by: Juneja, Ishank, et al.
Published: (2025)
Demystifying Online Clustering of Bandits: Enhanced Exploration Under Stochastic and Smoothed Adversarial Contexts
by: Li, Zhuohua, et al.
Published: (2025)
by: Li, Zhuohua, et al.
Published: (2025)
Unlearning Offline Stochastic Multi-Armed Bandits
by: Ye, Zichun, et al.
Published: (2026)
by: Ye, Zichun, et al.
Published: (2026)
Multi-Agent Stochastic Bandits Robust to Adversarial Corruptions
by: Ghaffari, Fatemeh, et al.
Published: (2024)
by: Ghaffari, Fatemeh, et al.
Published: (2024)
HiLoRA: Adaptive Hierarchical LoRA Routing for Training-Free Domain Generalization
by: Han, Ziyi, et al.
Published: (2025)
by: Han, Ziyi, et al.
Published: (2025)
Neural Bandit Based Optimal LLM Selection for a Pipeline of Subtasks
by: Atalar, Baran, et al.
Published: (2025)
by: Atalar, Baran, et al.
Published: (2025)
Near-Optimal Regret for Efficient Stochastic Combinatorial Semi-Bandits
by: Ye, Zichun, et al.
Published: (2025)
by: Ye, Zichun, et al.
Published: (2025)
Merit-based Fair Combinatorial Semi-Bandit with Unrestricted Feedback Delays
by: Chen, Ziqun, et al.
Published: (2024)
by: Chen, Ziqun, et al.
Published: (2024)
Federated Contextual Cascading Bandits with Asynchronous Communication and Heterogeneous Users
by: Yang, Hantao, et al.
Published: (2024)
by: Yang, Hantao, et al.
Published: (2024)
Memory-Based Advantage Shaping for LLM-Guided Reinforcement Learning
by: Nourzad, Narjes, et al.
Published: (2026)
by: Nourzad, Narjes, et al.
Published: (2026)
MIRA: Memory-Integrated Reinforcement Learning Agent with Limited LLM Guidance
by: Nourzad, Narjes, et al.
Published: (2026)
by: Nourzad, Narjes, et al.
Published: (2026)
M3Net: A Multi-Metric Mixture of Experts Network Digital Twin with Graph Neural Networks
by: Guda, Blessed, et al.
Published: (2025)
by: Guda, Blessed, et al.
Published: (2025)
Recommenadation aided Caching using Combinatorial Multi-armed Bandits
by: J, Pavamana K, et al.
Published: (2024)
by: J, Pavamana K, et al.
Published: (2024)
Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning
by: Chai, Jinhang, et al.
Published: (2025)
by: Chai, Jinhang, et al.
Published: (2025)
Similar Items
-
Batch-Size Independent Regret Bounds for Combinatorial Semi-Bandits with Probabilistically Triggered Arms or Independent Arms
by: Liu, Xutong, et al.
Published: (2022) -
Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
by: Liu, Xutong, et al.
Published: (2025) -
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
by: Liu, Xutong, et al.
Published: (2023) -
A Unified Online-Offline Framework for Co-Branding Campaign Recommendations
by: Dai, Xiangxiang, et al.
Published: (2025) -
Combinatorial Logistic Bandits
by: Liu, Xutong, et al.
Published: (2024)