Near-Optimal Online Deployment and Routing for Streaming LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Shaoang, Li, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915760908009472
author Li, Shaoang
Li, Jian
author_facet Li, Shaoang
Li, Jian
contents The rapid pace at which new large language models (LLMs) appear, and older ones become obsolete, forces providers to manage a streaming inventory under a strict concurrency cap and per-query cost budgets. We cast this as an online decision problem that couples stage-wise deployment (at fixed maintenance windows) with per-query routing among live models. We introduce StageRoute, a hierarchical algorithm that (i) optimistically selects up to $M_{\max}$ models for the next stage using reward upper-confidence and cost lower-confidence bounds, and (ii) routes each incoming query by solving a budget- and throughput-constrained bandit subproblem over the deployed set. We prove a regret of $\tilde{\mathcal{O}}(T^{2/3})$ with a matching lower bound, establishing near-optimality, and validate the theory empirically: StageRoute tracks a strong oracle under tight budgets across diverse workloads.
format Preprint
id arxiv_https___arxiv_org_abs_2506_17254
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Near-Optimal Online Deployment and Routing for Streaming LLMs
Li, Shaoang
Li, Jian
Machine Learning
Artificial Intelligence
The rapid pace at which new large language models (LLMs) appear, and older ones become obsolete, forces providers to manage a streaming inventory under a strict concurrency cap and per-query cost budgets. We cast this as an online decision problem that couples stage-wise deployment (at fixed maintenance windows) with per-query routing among live models. We introduce StageRoute, a hierarchical algorithm that (i) optimistically selects up to $M_{\max}$ models for the next stage using reward upper-confidence and cost lower-confidence bounds, and (ii) routes each incoming query by solving a budget- and throughput-constrained bandit subproblem over the deployed set. We prove a regret of $\tilde{\mathcal{O}}(T^{2/3})$ with a matching lower bound, establishing near-optimality, and validate the theory empirically: StageRoute tracks a strong oracle under tight budgets across diverse workloads.
title Near-Optimal Online Deployment and Routing for Streaming LLMs
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.17254