Near-Optimal Online Deployment and Routing for Streaming LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Shaoang, Li, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving
by: Li, Shaoang, et al.
Published: (2026)
by: Li, Shaoang, et al.
Published: (2026)
COPF: An Online Framework for Deployment-Stable Counterfactual Fairness in Evolving Graphs
by: Li, Sheng'en, et al.
Published: (2026)
by: Li, Sheng'en, et al.
Published: (2026)
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
Merging Embedded Topics with Optimal Transport for Online Topic Modeling on Data Streams
by: Granese, Federica, et al.
Published: (2025)
by: Granese, Federica, et al.
Published: (2025)
Constrained Feedback Learning for Non-Stationary Multi-Armed Bandits
by: Li, Shaoang, et al.
Published: (2025)
by: Li, Shaoang, et al.
Published: (2025)
Near-Optimal Partially Observable Reinforcement Learning with Partial Online State Information
by: Shi, Ming, et al.
Published: (2023)
by: Shi, Ming, et al.
Published: (2023)
Efficiently Deploying LLMs with Controlled Risk
by: Zellinger, Michael J., et al.
Published: (2024)
by: Zellinger, Michael J., et al.
Published: (2024)
RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs
by: Geuter, Jonathan, et al.
Published: (2025)
by: Geuter, Jonathan, et al.
Published: (2025)
StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs
by: Luo, Qijun, et al.
Published: (2025)
by: Luo, Qijun, et al.
Published: (2025)
RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs
by: Xu, Zhiyuan, et al.
Published: (2026)
by: Xu, Zhiyuan, et al.
Published: (2026)
Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices
by: Qin, Ruiyang, et al.
Published: (2024)
by: Qin, Ruiyang, et al.
Published: (2024)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning
by: Nilaksh, et al.
Published: (2026)
by: Nilaksh, et al.
Published: (2026)
Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
by: Chen, Yuanteng, et al.
Published: (2025)
by: Chen, Yuanteng, et al.
Published: (2025)
Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams
by: Liu, Zewen, et al.
Published: (2026)
by: Liu, Zewen, et al.
Published: (2026)
RouteLLM: Learning to Route LLMs with Preference Data
by: Ong, Isaac, et al.
Published: (2024)
by: Ong, Isaac, et al.
Published: (2024)
Neural Cluster First, Route Second: One-Shot Capacitated Vehicle Routing via Differentiable Optimal Transport
by: Chin, Samuel J. K., et al.
Published: (2026)
by: Chin, Samuel J. K., et al.
Published: (2026)
RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
by: Fernandez, Nigel, et al.
Published: (2025)
by: Fernandez, Nigel, et al.
Published: (2025)
Iterative Deployment Improves Planning Skills in LLMs
by: Corrêa, Augusto B., et al.
Published: (2025)
by: Corrêa, Augusto B., et al.
Published: (2025)
Online Sparse Feature Selection in Data Streams via Differential Evolution
by: Xu, Ruiyang
Published: (2025)
by: Xu, Ruiyang
Published: (2025)
Realizable Abstractions: Near-Optimal Hierarchical Reinforcement Learning
by: Cipollone, Roberto, et al.
Published: (2025)
by: Cipollone, Roberto, et al.
Published: (2025)
Invariant Causal Routing for Governing Social Norms in Online Market Economies
by: Yu, Xiangning, et al.
Published: (2026)
by: Yu, Xiangning, et al.
Published: (2026)
Retro-BLEU: Quantifying Chemical Plausibility of Retrosynthesis Routes through Reaction Template Sequence Analysis
by: Li, Junren, et al.
Published: (2023)
by: Li, Junren, et al.
Published: (2023)
OLR-WA: Online Weighted Average Linear Regression in Multivariate Data Streams
by: Abu-Shaira, Mohammad, et al.
Published: (2025)
by: Abu-Shaira, Mohammad, et al.
Published: (2025)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
by: Jiang, Yukun, et al.
Published: (2026)
by: Jiang, Yukun, et al.
Published: (2026)
Q-Palette: Fractional-Bit Quantizers Toward Optimal Bit Allocation for Efficient LLM Deployment
by: Lee, Deokjae, et al.
Published: (2025)
by: Lee, Deokjae, et al.
Published: (2025)
Learning to Route LLMs with Confidence Tokens
by: Chuang, Yu-Neng, et al.
Published: (2024)
by: Chuang, Yu-Neng, et al.
Published: (2024)
ICL-Router: In-Context Learned Model Representations for LLM Routing
by: Wang, Chenxu, et al.
Published: (2025)
by: Wang, Chenxu, et al.
Published: (2025)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Minimax Optimality and Spectral Routing for Majority-Vote Ensembles under Markov Dependence
by: Shihab, Ibne Farabi, et al.
Published: (2026)
by: Shihab, Ibne Farabi, et al.
Published: (2026)
Stream-level flow matching with Gaussian processes
by: Wei, Ganchao, et al.
Published: (2024)
by: Wei, Ganchao, et al.
Published: (2024)
A Survey of Route Recommendations: Methods, Applications, and Opportunities
by: Zhang, Shiming, et al.
Published: (2024)
by: Zhang, Shiming, et al.
Published: (2024)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
by: Zhang, Yiqi, et al.
Published: (2026)
by: Zhang, Yiqi, et al.
Published: (2026)
Near-Optimal Experiment Design in Linear non-Gaussian Cyclic Models
by: Sharifian, Ehsan, et al.
Published: (2025)
by: Sharifian, Ehsan, et al.
Published: (2025)
Near-Optimal Dynamic Matching via Coarsening with Application to Heart Transplantation
by: Zilberstein, Itai, et al.
Published: (2026)
by: Zilberstein, Itai, et al.
Published: (2026)
A Theory of Training Profit-Optimal LLMs
by: Hao, Sophie, et al.
Published: (2026)
by: Hao, Sophie, et al.
Published: (2026)
DTMM: Deploying TinyML Models on Extremely Weak IoT Devices with Pruning
by: Han, Lixiang, et al.
Published: (2024)
by: Han, Lixiang, et al.
Published: (2024)
Dr.LLM: Dynamic Layer Routing in LLMs
by: Heakl, Ahmed, et al.
Published: (2025)
by: Heakl, Ahmed, et al.
Published: (2025)
Fair Streaming Feature Selection
by: Duan, Zhangling, et al.
Published: (2024)
by: Duan, Zhangling, et al.
Published: (2024)
Extension OL-MDISF: Online Learning from Mix-Typed, Drifted, and Incomplete Streaming Features
by: Zhuo, Shengda, et al.
Published: (2025)
by: Zhuo, Shengda, et al.
Published: (2025)
Similar Items
-
POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving
by: Li, Shaoang, et al.
Published: (2026) -
COPF: An Online Framework for Deployment-Stable Counterfactual Fairness in Evolving Graphs
by: Li, Sheng'en, et al.
Published: (2026) -
Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing
by: Zhang, Zeyu, et al.
Published: (2026) -
Merging Embedded Topics with Optimal Transport for Online Topic Modeling on Data Streams
by: Granese, Federica, et al.
Published: (2025) -
Constrained Feedback Learning for Non-Stationary Multi-Armed Bandits
by: Li, Shaoang, et al.
Published: (2025)