LLM Serving Optimization with Variable Prefill and Decode Lengths
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Meixuan, Ye, Yinyu, Zhou, Zijie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025)
by: Chen, Zixi, et al.
Published: (2025)
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
by: Nie, Chengyi, et al.
Published: (2026)
by: Nie, Chengyi, et al.
Published: (2026)
Online Scheduling for LLM Inference with KV Cache Constraints
by: Jaillet, Patrick, et al.
Published: (2025)
by: Jaillet, Patrick, et al.
Published: (2025)
From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling
by: Lin, Jianghao, et al.
Published: (2026)
by: Lin, Jianghao, et al.
Published: (2026)
OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
by: Ao, Ruicheng, et al.
Published: (2026)
by: Ao, Ruicheng, et al.
Published: (2026)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling
by: Hao, Yongchang, et al.
Published: (2026)
by: Hao, Yongchang, et al.
Published: (2026)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
Achieving Instance-dependent Sample Complexity for Constrained Markov Decision Process
by: Jiang, Jiashuo, et al.
Published: (2024)
by: Jiang, Jiashuo, et al.
Published: (2024)
A Minimalist Bayesian Framework for Stochastic Optimization
by: Wang, Kaizheng
Published: (2025)
by: Wang, Kaizheng
Published: (2025)
Federated Instrumental Variable Analysis via Federated Generalized Method of Moments
by: Geetika, et al.
Published: (2025)
by: Geetika, et al.
Published: (2025)
Contextual Distributionally Robust Optimization with Causal and Continuous Structure: An Interpretable and Tractable Approach
by: Zhang, Fenglin, et al.
Published: (2026)
by: Zhang, Fenglin, et al.
Published: (2026)
Benchmarking PtO and PnO Methods in the Predictive Combinatorial Optimization Regime
by: Geng, Haoyu, et al.
Published: (2023)
by: Geng, Haoyu, et al.
Published: (2023)
ARO: A New Lens On Matrix Optimization For Large Models
by: Gong, Wenbo, et al.
Published: (2026)
by: Gong, Wenbo, et al.
Published: (2026)
Optimizing the Optimizer for Physics-Informed Neural Networks and Kolmogorov-Arnold Networks
by: Kiyani, Elham, et al.
Published: (2025)
by: Kiyani, Elham, et al.
Published: (2025)
Demystifying Manifold Constraints in LLM Pre-training
by: An, Kang, et al.
Published: (2026)
by: An, Kang, et al.
Published: (2026)
The Newton-Muon Optimizer
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Riemannian Bilevel Optimization
by: Dutta, Sanchayan, et al.
Published: (2024)
by: Dutta, Sanchayan, et al.
Published: (2024)
MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters
by: Sharifnassab, Arsalan, et al.
Published: (2024)
by: Sharifnassab, Arsalan, et al.
Published: (2024)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
by: Fernando, Heshan, et al.
Published: (2024)
by: Fernando, Heshan, et al.
Published: (2024)
From Large Language Models and Optimization to Decision Optimization CoPilot: A Research Manifesto
by: Wasserkrug, Segev, et al.
Published: (2024)
by: Wasserkrug, Segev, et al.
Published: (2024)
Client-Centric Federated Adaptive Optimization
by: Sun, Jianhui, et al.
Published: (2025)
by: Sun, Jianhui, et al.
Published: (2025)
On the Condition Number Dependency in Bilevel Optimization
by: Chen, Lesi, et al.
Published: (2025)
by: Chen, Lesi, et al.
Published: (2025)
Budget-aware Auto Optimizer Configurator
by: Liu, Kang, et al.
Published: (2026)
by: Liu, Kang, et al.
Published: (2026)
Jacobian Descent for Multi-Objective Optimization
by: Quinton, Pierre, et al.
Published: (2024)
by: Quinton, Pierre, et al.
Published: (2024)
Optimization and Generalization Guarantees for Weight Normalization
by: Cisneros-Velarde, Pedro, et al.
Published: (2024)
by: Cisneros-Velarde, Pedro, et al.
Published: (2024)
Neural Solver Selection for Combinatorial Optimization
by: Gao, Chengrui, et al.
Published: (2024)
by: Gao, Chengrui, et al.
Published: (2024)
Dynamic Memory Based Adaptive Optimization
by: Szegedy, Balázs, et al.
Published: (2024)
by: Szegedy, Balázs, et al.
Published: (2024)
Zeroth-Order Optimization Finds Flat Minima
by: Zhang, Liang, et al.
Published: (2025)
by: Zhang, Liang, et al.
Published: (2025)
Tree-Preconditioned Differentiable Optimization and Axioms as Layers
by: Liao, Yuexin
Published: (2025)
by: Liao, Yuexin
Published: (2025)
Search-Optimized Quantization in Biomedical Ontology Alignment
by: Bouaggad, Oussama, et al.
Published: (2025)
by: Bouaggad, Oussama, et al.
Published: (2025)
Understanding Optimization in Deep Learning with Central Flows
by: Cohen, Jeremy M., et al.
Published: (2024)
by: Cohen, Jeremy M., et al.
Published: (2024)
Hierarchical Mixture-of-Experts with Two-Stage Optimization
by: Molodtsov, Gleb, et al.
Published: (2026)
by: Molodtsov, Gleb, et al.
Published: (2026)
Bayesian Optimization for Hyperparameters Tuning in Neural Networks
by: Onorato, Gabriele
Published: (2024)
by: Onorato, Gabriele
Published: (2024)
Intersectional Fairness via Mixed-Integer Optimization
by: Němeček, Jiří, et al.
Published: (2026)
by: Němeček, Jiří, et al.
Published: (2026)
Anytime Training with Schedule-Free Spectral Optimization
by: Apte, Anuj, et al.
Published: (2026)
by: Apte, Anuj, et al.
Published: (2026)
Constructing Industrial-Scale Optimization Modeling Benchmark
by: Li, Zhong, et al.
Published: (2026)
by: Li, Zhong, et al.
Published: (2026)
LLM Embeddings Improve Test-time Adaptation to Tabular $Y|X$-Shifts
by: Zeng, Yibo, et al.
Published: (2024)
by: Zeng, Yibo, et al.
Published: (2024)
On Some Tunable Multi-fidelity Bayesian Optimization Frameworks
by: Manoj, Arjun, et al.
Published: (2025)
by: Manoj, Arjun, et al.
Published: (2025)
Similar Items
-
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025) -
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
by: Nie, Chengyi, et al.
Published: (2026) -
Online Scheduling for LLM Inference with KV Cache Constraints
by: Jaillet, Patrick, et al.
Published: (2025) -
From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling
by: Lin, Jianghao, et al.
Published: (2026) -
OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
by: Ao, Ruicheng, et al.
Published: (2026)