A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Nie, Chengyi, Si, Nian, Zhou, Zijie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Online Scheduling for LLM Inference with KV Cache Constraints
by: Jaillet, Patrick, et al.
Published: (2025)
by: Jaillet, Patrick, et al.
Published: (2025)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025)
by: Chen, Zixi, et al.
Published: (2025)
A Queueing-Theoretic Framework for Dynamic Attack Surfaces: Data-Integrated Risk Analysis and Adaptive Defense
by: Yun, Jihyeon, et al.
Published: (2026)
by: Yun, Jihyeon, et al.
Published: (2026)
LLM Serving Optimization with Variable Prefill and Decode Lengths
by: Wang, Meixuan, et al.
Published: (2025)
by: Wang, Meixuan, et al.
Published: (2025)
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
by: Ao, Ruicheng, et al.
Published: (2025)
by: Ao, Ruicheng, et al.
Published: (2025)
Demystifying Manifold Constraints in LLM Pre-training
by: An, Kang, et al.
Published: (2026)
by: An, Kang, et al.
Published: (2026)
From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling
by: Lin, Jianghao, et al.
Published: (2026)
by: Lin, Jianghao, et al.
Published: (2026)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
Stochastic Optimization with Constraints: A Non-asymptotic Instance-Dependent Analysis
by: Khamaru, Koulik
Published: (2024)
by: Khamaru, Koulik
Published: (2024)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
by: Han, X. Y., et al.
Published: (2025)
by: Han, X. Y., et al.
Published: (2025)
T-SKM-Net: Trainable Neural Network Framework for Linear Constraint Satisfaction via Sampling Kaczmarz-Motzkin Method
by: Zhu, Haoyu, et al.
Published: (2025)
by: Zhu, Haoyu, et al.
Published: (2025)
Towards Efficient Constraint Handling in Neural Solvers for Routing Problems
by: Bi, Jieyi, et al.
Published: (2026)
by: Bi, Jieyi, et al.
Published: (2026)
CLCR: Contrastive Learning-based Constraint Reordering for Efficient MILP Solving
by: Zeng, Shuli, et al.
Published: (2025)
by: Zeng, Shuli, et al.
Published: (2025)
Hierarchical Deep Reinforcement Learning Framework for Multi-Year Asset Management Under Budget Constraints
by: Fard, Amir, et al.
Published: (2025)
by: Fard, Amir, et al.
Published: (2025)
Theoretical and Empirical Advances in Forest Pruning
by: Dorador, Albert
Published: (2024)
by: Dorador, Albert
Published: (2024)
Optimism Stabilizes Thompson Sampling for Adaptive Inference
by: Yan, Shunxing, et al.
Published: (2026)
by: Yan, Shunxing, et al.
Published: (2026)
Dynamic Memory Based Adaptive Optimization
by: Szegedy, Balázs, et al.
Published: (2024)
by: Szegedy, Balázs, et al.
Published: (2024)
A Rod Flow Model for Adam at the Edge of Stability
by: Regis, Eric, et al.
Published: (2026)
by: Regis, Eric, et al.
Published: (2026)
New Hybrid Fine-Tuning Paradigm for LLMs: Algorithm Design and Convergence Analysis Framework
by: Ma, Shaocong, et al.
Published: (2026)
by: Ma, Shaocong, et al.
Published: (2026)
Stability of Transformers under Layer Normalization
by: Kan, Kelvin, et al.
Published: (2025)
by: Kan, Kelvin, et al.
Published: (2025)
How Memory in Optimization Algorithms Implicitly Modifies the Loss
by: Cattaneo, Matias D., et al.
Published: (2025)
by: Cattaneo, Matias D., et al.
Published: (2025)
Convex and Bilevel Optimization for Neuro-Symbolic Inference and Learning
by: Dickens, Charles, et al.
Published: (2024)
by: Dickens, Charles, et al.
Published: (2024)
Rod Flow: A Continuous-Time Model for Gradient Descent at the Edge of Stability
by: Regis, Eric, et al.
Published: (2026)
by: Regis, Eric, et al.
Published: (2026)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
by: Liao, Fangshuo, et al.
Published: (2026)
by: Liao, Fangshuo, et al.
Published: (2026)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
Active Inference for Energy Control and Planning in Smart Buildings and Communities
by: Nazemi, Seyyed Danial, et al.
Published: (2025)
by: Nazemi, Seyyed Danial, et al.
Published: (2025)
Constraint-Anchored Attribution: Feasibility-Certified Counterfactuals and Bonferroni-PAC Sufficient Subsets for Neural CO Policies
by: Lafifi, Sohaib
Published: (2026)
by: Lafifi, Sohaib
Published: (2026)
Queueing-Aware Optimization of Reasoning Tokens for Accuracy-Latency Trade-offs in LLM Servers
by: Ozbas, Emre, et al.
Published: (2026)
by: Ozbas, Emre, et al.
Published: (2026)
A Minimalist Bayesian Framework for Stochastic Optimization
by: Wang, Kaizheng
Published: (2025)
by: Wang, Kaizheng
Published: (2025)
Locally Interdependent Multi-Agent MDP: Theoretical Framework for Decentralized Agents with Dynamic Dependencies
by: DeWeese, Alex, et al.
Published: (2024)
by: DeWeese, Alex, et al.
Published: (2024)
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
by: Xiao, Nachuan, et al.
Published: (2023)
by: Xiao, Nachuan, et al.
Published: (2023)
Online Submodular Maximization via Online Convex Optimization
by: Salem, Tareq Si, et al.
Published: (2023)
by: Salem, Tareq Si, et al.
Published: (2023)
A Unified Framework for Gradient Aggregation in Multi-Objective Optimization
by: Hu, Zeou, et al.
Published: (2026)
by: Hu, Zeou, et al.
Published: (2026)
Stabilizing reinforcement learning control: A modular framework for optimizing over all stable behavior
by: Lawrence, Nathan P., et al.
Published: (2023)
by: Lawrence, Nathan P., et al.
Published: (2023)
MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters
by: Sharifnassab, Arsalan, et al.
Published: (2024)
by: Sharifnassab, Arsalan, et al.
Published: (2024)
A Retention-Centric Framework for Continual Learning with Guaranteed Model Developmental Safety
by: Li, Gang, et al.
Published: (2024)
by: Li, Gang, et al.
Published: (2024)
LLM Embeddings Improve Test-time Adaptation to Tabular $Y|X$-Shifts
by: Zeng, Yibo, et al.
Published: (2024)
by: Zeng, Yibo, et al.
Published: (2024)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Stability of Primal-Dual Gradient Flow Dynamics for Multi-Block Convex Optimization Problems
by: Ozaslan, Ibrahim K., et al.
Published: (2024)
by: Ozaslan, Ibrahim K., et al.
Published: (2024)
On Some Tunable Multi-fidelity Bayesian Optimization Frameworks
by: Manoj, Arjun, et al.
Published: (2025)
by: Manoj, Arjun, et al.
Published: (2025)
Similar Items
-
Online Scheduling for LLM Inference with KV Cache Constraints
by: Jaillet, Patrick, et al.
Published: (2025) -
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025) -
A Queueing-Theoretic Framework for Dynamic Attack Surfaces: Data-Integrated Risk Analysis and Adaptive Defense
by: Yun, Jihyeon, et al.
Published: (2026) -
LLM Serving Optimization with Variable Prefill and Decode Lengths
by: Wang, Meixuan, et al.
Published: (2025) -
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
by: Ao, Ruicheng, et al.
Published: (2025)