Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Ao, Ruicheng, Luo, Gan, Simchi-Levi, David, Wang, Xinshang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OptPipe: Memory- and Scheduling-Optimized Pipeline Parallelism for LLM Training
by: Li, Hongpei, et al.
Published: (2025)
by: Li, Hongpei, et al.
Published: (2025)
OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
by: Ao, Ruicheng, et al.
Published: (2026)
by: Ao, Ruicheng, et al.
Published: (2026)
Constraint Programming Models For Serial Batch Scheduling With Minimum Batch Size
by: Huertas, Jorge A., et al.
Published: (2025)
by: Huertas, Jorge A., et al.
Published: (2025)
ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research
by: Ao, Ruicheng, et al.
Published: (2026)
by: Ao, Ruicheng, et al.
Published: (2026)
FADAS: Towards Federated Adaptive Asynchronous Optimization
by: Wang, Yujia, et al.
Published: (2024)
by: Wang, Yujia, et al.
Published: (2024)
Correlated Quantization for Faster Nonconvex Distributed Optimization
by: Panferov, Andrei, et al.
Published: (2024)
by: Panferov, Andrei, et al.
Published: (2024)
A Communication and Computation Efficient Fully First-order Method for Decentralized Bilevel Optimization
by: Wen, Min, et al.
Published: (2024)
by: Wen, Min, et al.
Published: (2024)
Several Performance Bounds on Decentralized Online Optimization are Highly Conservative and Potentially Misleading
by: Meunier, Erwan, et al.
Published: (2025)
by: Meunier, Erwan, et al.
Published: (2025)
CONGO: Compressive Online Gradient Optimization
by: Carleton, Jeremy, et al.
Published: (2024)
by: Carleton, Jeremy, et al.
Published: (2024)
Distributed Difference of Convex Optimization
by: Khatana, Vivek, et al.
Published: (2024)
by: Khatana, Vivek, et al.
Published: (2024)
A Double Tracking Method for Optimization with Decentralized Generalized Orthogonality Constraints
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
MAST: Model-Agnostic Sparsified Training
by: Demidovich, Yury, et al.
Published: (2023)
by: Demidovich, Yury, et al.
Published: (2023)
Byzantine Robustness and Partial Participation Can Be Achieved at Once: Just Clip Gradient Differences
by: Malinovsky, Grigory, et al.
Published: (2023)
by: Malinovsky, Grigory, et al.
Published: (2023)
Communication Efficient Distributed Training with Distributed Lion
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
Distributed Constraint-Coupled Optimization: Harnessing ADMM-consensus for robustness
by: Messilem, Mohamed Abdelmouamin, et al.
Published: (2025)
by: Messilem, Mohamed Abdelmouamin, et al.
Published: (2025)
Decentralized Nonsmooth Nonconvex Optimization with Client Sampling
by: Chen, Xinyan, et al.
Published: (2026)
by: Chen, Xinyan, et al.
Published: (2026)
A Hybrid Stochastic Gradient Tracking Method for Distributed Online Optimization Over Time-Varying Directed Networks
by: Shi, Xinli, et al.
Published: (2025)
by: Shi, Xinli, et al.
Published: (2025)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
by: Yu, Jiahuan, et al.
Published: (2026)
by: Yu, Jiahuan, et al.
Published: (2026)
Large-Scale LLM Inference with Heterogeneous Workloads: Prefill-Decode Contention and Asymptotically Optimal Control
by: Lin, Ruihan, et al.
Published: (2026)
by: Lin, Ruihan, et al.
Published: (2026)
Decentralized Gradient-Free Methods for Stochastic Non-Smooth Non-Convex Optimization
by: Lin, Zhenwei, et al.
Published: (2023)
by: Lin, Zhenwei, et al.
Published: (2023)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
by: Jiang, Xuanlin, et al.
Published: (2024)
by: Jiang, Xuanlin, et al.
Published: (2024)
Optimality in Decentralized Optimization under Bandwidth Constraints
by: Tyurin, Alexander
Published: (2026)
by: Tyurin, Alexander
Published: (2026)
Online Distributed Learning with Quantized Finite-Time Coordination
by: Bastianello, Nicola, et al.
Published: (2023)
by: Bastianello, Nicola, et al.
Published: (2023)
Towards Dynamic Resource Allocation and Client Scheduling in Hierarchical Federated Learning: A Two-Phase Deep Reinforcement Learning Approach
by: Chen, Xiaojing, et al.
Published: (2024)
by: Chen, Xiaojing, et al.
Published: (2024)
Efficient Adaptive Federated Optimization
by: Lee, Su Hyeong, et al.
Published: (2024)
by: Lee, Su Hyeong, et al.
Published: (2024)
Resource Allocation for Stable LLM Training in Mobile Edge Computing
by: Liu, Chang, et al.
Published: (2024)
by: Liu, Chang, et al.
Published: (2024)
Rack Position Optimization in Large-Scale Heterogeneous Data Centers
by: Chen, Chang-Lin, et al.
Published: (2025)
by: Chen, Chang-Lin, et al.
Published: (2025)
On Principled Local Optimization Methods for Federated Learning
by: Yuan, Honglin
Published: (2024)
by: Yuan, Honglin
Published: (2024)
Optimizing Stochastic Gradient Push under Broadcast Communications
by: Nguyen, Tuan, et al.
Published: (2026)
by: Nguyen, Tuan, et al.
Published: (2026)
A Single-Loop Algorithm for Decentralized Bilevel Optimization
by: Dong, Youran, et al.
Published: (2023)
by: Dong, Youran, et al.
Published: (2023)
Convergence of Sign-based Random Reshuffling Algorithms for Nonconvex Optimization
by: Qin, Zhen, et al.
Published: (2023)
by: Qin, Zhen, et al.
Published: (2023)
Communication-Efficient Federated Optimization over Semi-Decentralized Networks
by: Wang, He, et al.
Published: (2023)
by: Wang, He, et al.
Published: (2023)
A Stochastic Approximation Approach for Efficient Decentralized Optimization on Random Networks
by: Yau, Chung-Yiu, et al.
Published: (2024)
by: Yau, Chung-Yiu, et al.
Published: (2024)
Achieving Near-Optimal Convergence for Distributed Minimax Optimization with Adaptive Stepsizes
by: Huang, Yan, et al.
Published: (2024)
by: Huang, Yan, et al.
Published: (2024)
Accelerating Distributed Optimization: A Primal-Dual Perspective on Local Steps
by: Yang, Junchi, et al.
Published: (2024)
by: Yang, Junchi, et al.
Published: (2024)
Lower Bounds and Accelerated Algorithms in Distributed Stochastic Optimization with Communication Compression
by: He, Yutong, et al.
Published: (2023)
by: He, Yutong, et al.
Published: (2023)
Unbiased Compression Saves Communication in Distributed Optimization: When and How Much?
by: He, Yutong, et al.
Published: (2023)
by: He, Yutong, et al.
Published: (2023)
A Graph-Based, Distributed Memory, Modeling Abstraction for Optimization
by: Cole, David L., et al.
Published: (2025)
by: Cole, David L., et al.
Published: (2025)
Lodestar: An Online-Learning LLM Inference Router
by: Lim, Gangmuk, et al.
Published: (2026)
by: Lim, Gangmuk, et al.
Published: (2026)
Accelerated Methods with Compressed Communications for Distributed Optimization Problems under Data Similarity
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Similar Items
-
OptPipe: Memory- and Scheduling-Optimized Pipeline Parallelism for LLM Training
by: Li, Hongpei, et al.
Published: (2025) -
OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
by: Ao, Ruicheng, et al.
Published: (2026) -
Constraint Programming Models For Serial Batch Scheduling With Minimum Batch Size
by: Huertas, Jorge A., et al.
Published: (2025) -
ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research
by: Ao, Ruicheng, et al.
Published: (2026) -
FADAS: Towards Federated Adaptive Asynchronous Optimization
by: Wang, Yujia, et al.
Published: (2024)