Lodestar: An Online-Learning LLM Inference Router
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lim, Gangmuk, Zhao, Wanyu, Godfrey, Brighten, Shan, Jiaxin, Xu, Le, Xie, Liguang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
par: Chen, Le, et autres
Publié: (2025)
par: Chen, Le, et autres
Publié: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
par: Zhu, Kan, et autres
Publié: (2025)
par: Zhu, Kan, et autres
Publié: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
par: Jiang, Xuanlin, et autres
Publié: (2024)
par: Jiang, Xuanlin, et autres
Publié: (2024)
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
par: The AIBrix Team, et autres
Publié: (2025)
par: The AIBrix Team, et autres
Publié: (2025)
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
par: Feng, Yicheng, et autres
Publié: (2026)
par: Feng, Yicheng, et autres
Publié: (2026)
Frontier: Simulating the Next Generation of LLM Inference Systems
par: Feng, Yicheng, et autres
Publié: (2025)
par: Feng, Yicheng, et autres
Publié: (2025)
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
par: Gond, Raja, et autres
Publié: (2026)
par: Gond, Raja, et autres
Publié: (2026)
Online Parallel Multi-Task Relationship Learning via Alternating Direction Method of Multipliers
par: Li, Ruiyu, et autres
Publié: (2024)
par: Li, Ruiyu, et autres
Publié: (2024)
An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
par: Dong, Hang, et autres
Publié: (2024)
par: Dong, Hang, et autres
Publié: (2024)
Niyama : Breaking the Silos of LLM Inference Serving
par: Goel, Kanishk, et autres
Publié: (2025)
par: Goel, Kanishk, et autres
Publié: (2025)
On Evaluating Performance of LLM Inference Serving Systems
par: Agrawal, Amey, et autres
Publié: (2025)
par: Agrawal, Amey, et autres
Publié: (2025)
MatKV: Trading Compute for Flash Storage in LLM Inference
par: Shin, Kun-Woo, et autres
Publié: (2025)
par: Shin, Kun-Woo, et autres
Publié: (2025)
Application of Machine Learning Optimization in Cloud Computing Resource Scheduling and Management
par: Zhang, Yifan, et autres
Publié: (2024)
par: Zhang, Yifan, et autres
Publié: (2024)
Federated Learning for Collaborative Inference Systems: The Case of Early Exit Networks
par: Kaplan, Caelin, et autres
Publié: (2024)
par: Kaplan, Caelin, et autres
Publié: (2024)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
par: Ye, Zihao, et autres
Publié: (2025)
par: Ye, Zihao, et autres
Publié: (2025)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
par: Jaiswal, Shashwat, et autres
Publié: (2025)
par: Jaiswal, Shashwat, et autres
Publié: (2025)
Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments
par: Yang, Yuzhe, et autres
Publié: (2024)
par: Yang, Yuzhe, et autres
Publié: (2024)
Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity
par: Yu, Donglin
Publié: (2026)
par: Yu, Donglin
Publié: (2026)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
par: Deshmukh, Dhruv, et autres
Publié: (2025)
par: Deshmukh, Dhruv, et autres
Publié: (2025)
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
par: Ao, Ruicheng, et autres
Publié: (2025)
par: Ao, Ruicheng, et autres
Publié: (2025)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
par: Deng, Xiumei, et autres
Publié: (2025)
par: Deng, Xiumei, et autres
Publié: (2025)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
par: Yu, Jiahuan, et autres
Publié: (2026)
par: Yu, Jiahuan, et autres
Publié: (2026)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
par: Rhee, Myunghyun, et autres
Publié: (2025)
par: Rhee, Myunghyun, et autres
Publié: (2025)
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
par: Levine, Reese, et autres
Publié: (2026)
par: Levine, Reese, et autres
Publié: (2026)
Context Parallelism for Scalable Million-Token Inference
par: Yang, Amy, et autres
Publié: (2024)
par: Yang, Amy, et autres
Publié: (2024)
Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
par: Zhang, Biyao, et autres
Publié: (2025)
par: Zhang, Biyao, et autres
Publié: (2025)
Convergence Analysis of Federated Learning Methods Using Backward Error Analysis
par: Lim, Jinwoo, et autres
Publié: (2025)
par: Lim, Jinwoo, et autres
Publié: (2025)
Adaptive Active Inference Agents for Heterogeneous and Lifelong Federated Learning
par: Danilenka, Anastasiya, et autres
Publié: (2024)
par: Danilenka, Anastasiya, et autres
Publié: (2024)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
par: Chen, Yanxi, et autres
Publié: (2023)
par: Chen, Yanxi, et autres
Publié: (2023)
EdgeRL: Reinforcement Learning-driven Deep Learning Model Inference Optimization at Edge
par: Mounesan, Motahare, et autres
Publié: (2024)
par: Mounesan, Motahare, et autres
Publié: (2024)
Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization
par: Li, Yang, et autres
Publié: (2025)
par: Li, Yang, et autres
Publié: (2025)
Online Client Scheduling and Resource Allocation for Efficient Federated Edge Learning
par: Gao, Zhidong, et autres
Publié: (2024)
par: Gao, Zhidong, et autres
Publié: (2024)
Context-Aware Inference via Performance Forecasting in Decentralized Learning Networks
par: Pfeffer, Joel, et autres
Publié: (2025)
par: Pfeffer, Joel, et autres
Publié: (2025)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
par: Zhao, Juntao, et autres
Publié: (2024)
par: Zhao, Juntao, et autres
Publié: (2024)
Efficient Serving of LLM Applications with Probabilistic Demand Modeling
par: Liu, Yifei, et autres
Publié: (2025)
par: Liu, Yifei, et autres
Publié: (2025)
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
par: Qin, Ruoyu, et autres
Publié: (2025)
par: Qin, Ruoyu, et autres
Publié: (2025)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
par: Huang, Shaoyuan, et autres
Publié: (2026)
par: Huang, Shaoyuan, et autres
Publié: (2026)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
par: Oakley, Joe, et autres
Publié: (2024)
par: Oakley, Joe, et autres
Publié: (2024)
DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge Devices
par: Peng, Xian, et autres
Publié: (2025)
par: Peng, Xian, et autres
Publié: (2025)
Documents similaires
-
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
par: Chen, Le, et autres
Publié: (2025) -
PolyServe: Efficient Multi-SLO Serving at Scale
par: Zhu, Kan, et autres
Publié: (2025) -
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
par: Jiang, Xuanlin, et autres
Publié: (2024) -
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
par: The AIBrix Team, et autres
Publié: (2025) -
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
par: Feng, Yicheng, et autres
Publié: (2026)