Online Multi-LLM Selection via Contextual Bandits under Unstructured Context Evolution
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Poon, Manhin, Dai, XiangXiang, Liu, Xutong, Kong, Fang, Lui, John C. S., Zuo, Jinhang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
von: Liu, Xutong, et al.
Veröffentlicht: (2023)
von: Liu, Xutong, et al.
Veröffentlicht: (2023)
Stochastic Bandits Robust to Adversarial Attacks
von: Wang, Xuchuang, et al.
Veröffentlicht: (2024)
von: Wang, Xuchuang, et al.
Veröffentlicht: (2024)
A Unified Online-Offline Framework for Co-Branding Campaign Recommendations
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2025)
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2025)
Offline Learning for Combinatorial Multi-armed Bandits
von: Liu, Xutong, et al.
Veröffentlicht: (2025)
von: Liu, Xutong, et al.
Veröffentlicht: (2025)
Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback
von: Zeng, Qirun, et al.
Veröffentlicht: (2026)
von: Zeng, Qirun, et al.
Veröffentlicht: (2026)
Cost-Effective Online Multi-LLM Selection with Versatile Reward Models
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2024)
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2024)
Fusing Reward and Dueling Feedback in Stochastic Bandits
von: Wang, Xuchuang, et al.
Veröffentlicht: (2025)
von: Wang, Xuchuang, et al.
Veröffentlicht: (2025)
Batch-Size Independent Regret Bounds for Combinatorial Semi-Bandits with Probabilistically Triggered Arms or Independent Arms
von: Liu, Xutong, et al.
Veröffentlicht: (2022)
von: Liu, Xutong, et al.
Veröffentlicht: (2022)
Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
von: Liu, Xutong, et al.
Veröffentlicht: (2025)
von: Liu, Xutong, et al.
Veröffentlicht: (2025)
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond
von: Liu, Xutong, et al.
Veröffentlicht: (2024)
von: Liu, Xutong, et al.
Veröffentlicht: (2024)
Combinatorial Logistic Bandits
von: Liu, Xutong, et al.
Veröffentlicht: (2024)
von: Liu, Xutong, et al.
Veröffentlicht: (2024)
Leveraging the Power of Conversations: Optimal Key Term Selection in Conversational Contextual Bandits
von: Liu, Maoli, et al.
Veröffentlicht: (2025)
von: Liu, Maoli, et al.
Veröffentlicht: (2025)
Federated Contextual Cascading Bandits with Asynchronous Communication and Heterogeneous Users
von: Yang, Hantao, et al.
Veröffentlicht: (2024)
von: Yang, Hantao, et al.
Veröffentlicht: (2024)
A Multi-Agent Conversational Bandit Approach to Online Evaluation and Selection of User-Aligned LLM Responses
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2025)
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2025)
Demystifying Online Clustering of Bandits: Enhanced Exploration Under Stochastic and Smoothed Adversarial Contexts
von: Li, Zhuohua, et al.
Veröffentlicht: (2025)
von: Li, Zhuohua, et al.
Veröffentlicht: (2025)
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
von: Zhang, Zeyu, et al.
Veröffentlicht: (2026)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2026)
Scaling Federated Linear Contextual Bandits via Sketching
von: Yang, Hantao, et al.
Veröffentlicht: (2026)
von: Yang, Hantao, et al.
Veröffentlicht: (2026)
Online Clustering of Dueling Bandits
von: Wang, Zhiyong, et al.
Veröffentlicht: (2025)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2025)
Offline Clustering of Linear Bandits: The Power of Clusters under Limited Data
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Jingyuan, et al.
Veröffentlicht: (2025)
Faster, Smaller, and Smarter: Task-Aware Expert Merging for Online MoE Inference
von: Han, Ziyi, et al.
Veröffentlicht: (2025)
von: Han, Ziyi, et al.
Veröffentlicht: (2025)
Multi-Agent Stochastic Bandits Robust to Adversarial Corruptions
von: Ghaffari, Fatemeh, et al.
Veröffentlicht: (2024)
von: Ghaffari, Fatemeh, et al.
Veröffentlicht: (2024)
Continuous Semantic Caching for Low-Cost LLM Serving
von: Atalar, Baran, et al.
Veröffentlicht: (2026)
von: Atalar, Baran, et al.
Veröffentlicht: (2026)
Quantum Algorithm for Online Exp-concave Optimization
von: He, Jianhao, et al.
Veröffentlicht: (2024)
von: He, Jianhao, et al.
Veröffentlicht: (2024)
Heterogeneous Multi-Agent Bandits with Parsimonious Hints
von: Mirfakhar, Amirmahdi, et al.
Veröffentlicht: (2025)
von: Mirfakhar, Amirmahdi, et al.
Veröffentlicht: (2025)
Online Learning to Rank under Corruption: A Robust Cascading Bandits Approach
von: Ghaffari, Fatemeh, et al.
Veröffentlicht: (2025)
von: Ghaffari, Fatemeh, et al.
Veröffentlicht: (2025)
AxiomVision: Accuracy-Guaranteed Adaptive Visual Model Selection for Perspective-Aware Video Analytics
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2024)
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2024)
Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection
von: Zeng, Qirun, et al.
Veröffentlicht: (2025)
von: Zeng, Qirun, et al.
Veröffentlicht: (2025)
Which LLM to Play? Convergence-Aware Online Model Selection with Time-Increasing Bandits
von: Xia, Yu, et al.
Veröffentlicht: (2024)
von: Xia, Yu, et al.
Veröffentlicht: (2024)
Active Context Selection Improves Simple Regret in Contextual Bandits
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2026)
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2026)
Large Language Model-Enhanced Multi-Armed Bandits
von: Sun, Jiahang, et al.
Veröffentlicht: (2025)
von: Sun, Jiahang, et al.
Veröffentlicht: (2025)
Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits
von: Pershin, Maksim, et al.
Veröffentlicht: (2026)
von: Pershin, Maksim, et al.
Veröffentlicht: (2026)
HiLoRA: Adaptive Hierarchical LoRA Routing for Training-Free Domain Generalization
von: Han, Ziyi, et al.
Veröffentlicht: (2025)
von: Han, Ziyi, et al.
Veröffentlicht: (2025)
Contextual Bandits for Unbounded Context Distributions
von: Zhao, Puning, et al.
Veröffentlicht: (2024)
von: Zhao, Puning, et al.
Veröffentlicht: (2024)
Causal Contextual Bandits with Adaptive Context
von: Madhavan, Rahul, et al.
Veröffentlicht: (2024)
von: Madhavan, Rahul, et al.
Veröffentlicht: (2024)
Constrained Contextual Bandits with Adversarial Contexts
von: Sarkar, Dhruv, et al.
Veröffentlicht: (2026)
von: Sarkar, Dhruv, et al.
Veröffentlicht: (2026)
Online Statistical Inference for Contextual Bandits via Stochastic Gradient Descent
von: Chang, Xiangyu, et al.
Veröffentlicht: (2022)
von: Chang, Xiangyu, et al.
Veröffentlicht: (2022)
Truthful Reverse Auctions for Adaptive Selection via Contextual Multi-Armed Bandits
von: Patra, Pronoy, et al.
Veröffentlicht: (2026)
von: Patra, Pronoy, et al.
Veröffentlicht: (2026)
Unlearning Offline Stochastic Multi-Armed Bandits
von: Ye, Zichun, et al.
Veröffentlicht: (2026)
von: Ye, Zichun, et al.
Veröffentlicht: (2026)
FedConPE: Efficient Federated Conversational Bandits with Heterogeneous Clients
von: Li, Zhuohua, et al.
Veröffentlicht: (2024)
von: Li, Zhuohua, et al.
Veröffentlicht: (2024)
Heterogeneous Multi-agent Multi-armed Bandits on Stochastic Block Models
von: Xu, Mengfan, et al.
Veröffentlicht: (2025)
von: Xu, Mengfan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Contextual Combinatorial Bandits with Probabilistically Triggered Arms
von: Liu, Xutong, et al.
Veröffentlicht: (2023) -
Stochastic Bandits Robust to Adversarial Attacks
von: Wang, Xuchuang, et al.
Veröffentlicht: (2024) -
A Unified Online-Offline Framework for Co-Branding Campaign Recommendations
von: Dai, Xiangxiang, et al.
Veröffentlicht: (2025) -
Offline Learning for Combinatorial Multi-armed Bandits
von: Liu, Xutong, et al.
Veröffentlicht: (2025) -
Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback
von: Zeng, Qirun, et al.
Veröffentlicht: (2026)