Online Scheduling for LLM Inference with KV Cache Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Jaillet, Patrick, Jiang, Jiashuo, Mellou, Konstantina, Molinaro, Marco, Podimata, Chara, Zhou, Zijie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
by: Nie, Chengyi, et al.
Published: (2026)
by: Nie, Chengyi, et al.
Published: (2026)
When Should you Offer an Upgrade: Online Upgrading Mechanisms for Resource Allocation
by: Jaillet, Patrick, et al.
Published: (2024)
by: Jaillet, Patrick, et al.
Published: (2024)
Grace Period is All You Need: Individual Fairness without Revenue Loss in Revenue Management
by: Jaillet, Patrick, et al.
Published: (2024)
by: Jaillet, Patrick, et al.
Published: (2024)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025)
by: Chen, Zixi, et al.
Published: (2025)
Online Resource Allocation with Convex-set Machine-Learned Advice
by: Golrezaei, Negin, et al.
Published: (2023)
by: Golrezaei, Negin, et al.
Published: (2023)
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
by: Ao, Ruicheng, et al.
Published: (2025)
by: Ao, Ruicheng, et al.
Published: (2025)
LLM Serving Optimization with Variable Prefill and Decode Lengths
by: Wang, Meixuan, et al.
Published: (2025)
by: Wang, Meixuan, et al.
Published: (2025)
LLM Embeddings Improve Test-time Adaptation to Tabular $Y|X$-Shifts
by: Zeng, Yibo, et al.
Published: (2024)
by: Zeng, Yibo, et al.
Published: (2024)
Integrated Offline and Online Learning to Solve a Large Class of Scheduling Problems
by: Liu, Anbang, et al.
Published: (2025)
by: Liu, Anbang, et al.
Published: (2025)
Demystifying Manifold Constraints in LLM Pre-training
by: An, Kang, et al.
Published: (2026)
by: An, Kang, et al.
Published: (2026)
Multi-Timescale Primal Dual Hybrid Gradient with Application to Distributed Optimization
by: Zhang, Junhui, et al.
Published: (2025)
by: Zhang, Junhui, et al.
Published: (2025)
The Road Less Scheduled
by: Defazio, Aaron, et al.
Published: (2024)
by: Defazio, Aaron, et al.
Published: (2024)
ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule
by: Huang, Yilie, et al.
Published: (2026)
by: Huang, Yilie, et al.
Published: (2026)
Solving Integrated Process Planning and Scheduling Problem via Graph Neural Network Based Deep Reinforcement Learning
by: Li, Hongpei, et al.
Published: (2024)
by: Li, Hongpei, et al.
Published: (2024)
Anytime Training with Schedule-Free Spectral Optimization
by: Apte, Anuj, et al.
Published: (2026)
by: Apte, Anuj, et al.
Published: (2026)
From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling
by: Lin, Jianghao, et al.
Published: (2026)
by: Lin, Jianghao, et al.
Published: (2026)
Towards Efficient Constraint Handling in Neural Solvers for Routing Problems
by: Bi, Jieyi, et al.
Published: (2026)
by: Bi, Jieyi, et al.
Published: (2026)
CLCR: Contrastive Learning-based Constraint Reordering for Efficient MILP Solving
by: Zeng, Shuli, et al.
Published: (2025)
by: Zeng, Shuli, et al.
Published: (2025)
Graph Neural Networks for the Offline Nanosatellite Task Scheduling Problem
by: Pacheco, Bruno Machado, et al.
Published: (2023)
by: Pacheco, Bruno Machado, et al.
Published: (2023)
Online Submodular Maximization via Online Convex Optimization
by: Salem, Tareq Si, et al.
Published: (2023)
by: Salem, Tareq Si, et al.
Published: (2023)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
Neural Combinatorial Optimization for Stochastic Flexible Job Shop Scheduling Problems
by: Smit, Igor G., et al.
Published: (2024)
by: Smit, Igor G., et al.
Published: (2024)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
by: Song, Minhak, et al.
Published: (2025)
by: Song, Minhak, et al.
Published: (2025)
Learning-Guided Rolling Horizon Optimization for Long-Horizon Flexible Job-Shop Scheduling
by: Li, Sirui, et al.
Published: (2025)
by: Li, Sirui, et al.
Published: (2025)
Learning to Price with Resource Constraints: From Full Information to Machine-Learned Prices
by: Ao, Ruicheng, et al.
Published: (2025)
by: Ao, Ruicheng, et al.
Published: (2025)
Stochastic Optimization with Constraints: A Non-asymptotic Instance-Dependent Analysis
by: Khamaru, Koulik
Published: (2024)
by: Khamaru, Koulik
Published: (2024)
Convex and Bilevel Optimization for Neuro-Symbolic Inference and Learning
by: Dickens, Charles, et al.
Published: (2024)
by: Dickens, Charles, et al.
Published: (2024)
Active Inference for Energy Control and Planning in Smart Buildings and Communities
by: Nazemi, Seyyed Danial, et al.
Published: (2025)
by: Nazemi, Seyyed Danial, et al.
Published: (2025)
BAGEL: Projection-Free Algorithm for Adversarially Constrained Online Convex Optimization
by: Lu, Yiyang, et al.
Published: (2025)
by: Lu, Yiyang, et al.
Published: (2025)
Constraint-Anchored Attribution: Feasibility-Certified Counterfactuals and Bonferroni-PAC Sufficient Subsets for Neural CO Policies
by: Lafifi, Sohaib
Published: (2026)
by: Lafifi, Sohaib
Published: (2026)
T-SKM-Net: Trainable Neural Network Framework for Linear Constraint Satisfaction via Sampling Kaczmarz-Motzkin Method
by: Zhu, Haoyu, et al.
Published: (2025)
by: Zhu, Haoyu, et al.
Published: (2025)
Learning Concave Bid Shading Strategies in Online Auctions via Measure-valued Proximal Optimization
by: Nodozi, Iman, et al.
Published: (2025)
by: Nodozi, Iman, et al.
Published: (2025)
Hierarchical Deep Reinforcement Learning Framework for Multi-Year Asset Management Under Budget Constraints
by: Fard, Amir, et al.
Published: (2025)
by: Fard, Amir, et al.
Published: (2025)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
Active Constraint Learning in High Dimensions from Demonstrations
by: Qiu, Zheng, et al.
Published: (2025)
by: Qiu, Zheng, et al.
Published: (2025)
OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
by: Ao, Ruicheng, et al.
Published: (2026)
by: Ao, Ruicheng, et al.
Published: (2026)
Achieving Instance-dependent Sample Complexity for Constrained Markov Decision Process
by: Jiang, Jiashuo, et al.
Published: (2024)
by: Jiang, Jiashuo, et al.
Published: (2024)
Clip-and-Verify: Linear Constraint-Driven Domain Clipping for Accelerating Neural Network Verification
by: Zhou, Duo, et al.
Published: (2025)
by: Zhou, Duo, et al.
Published: (2025)
Unveiling Hidden Pivotal Players with GoalNet: A GNN-Based Soccer Player Evaluation System
by: Jiang, Jacky Hao, et al.
Published: (2025)
by: Jiang, Jacky Hao, et al.
Published: (2025)
Similar Items
-
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
by: Nie, Chengyi, et al.
Published: (2026) -
When Should you Offer an Upgrade: Online Upgrading Mechanisms for Resource Allocation
by: Jaillet, Patrick, et al.
Published: (2024) -
Grace Period is All You Need: Individual Fairness without Revenue Loss in Revenue Management
by: Jaillet, Patrick, et al.
Published: (2024) -
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025) -
Online Resource Allocation with Convex-set Machine-Learned Advice
by: Golrezaei, Negin, et al.
Published: (2023)