PRISM: Fast Online LLM Serving via Scheduling-Memory Co-design
Fuente:
arXiv
Saved in:
| Main Authors: | Qu, Xingyu, Lin, Tianhao, Li, Yiqi, Chen, Zhiyu, Wang, Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Echo: Efficient Co-Scheduling of Hybrid Online-Offline Tasks for Large Language Model Serving
by: Wang, Zhibin, et al.
Published: (2025)
by: Wang, Zhibin, et al.
Published: (2025)
Locality-aware Fair Scheduling in LLM Serving
by: Cao, Shiyi, et al.
Published: (2025)
by: Cao, Shiyi, et al.
Published: (2025)
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
by: Gao, Shihong, et al.
Published: (2025)
by: Gao, Shihong, et al.
Published: (2025)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
by: Qiao, Yifan, et al.
Published: (2024)
by: Qiao, Yifan, et al.
Published: (2024)
PRISM: Parallel Residual Iterative Sequence Model
by: Jiang, Jie, et al.
Published: (2026)
by: Jiang, Jie, et al.
Published: (2026)
HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
by: Sun, Ting, et al.
Published: (2025)
by: Sun, Ting, et al.
Published: (2025)
TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling
by: Chen, Junyi, et al.
Published: (2025)
by: Chen, Junyi, et al.
Published: (2025)
FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
Preble: Efficient Distributed Prompt Scheduling for LLM Serving
by: Srivatsa, Vikranth, et al.
Published: (2024)
by: Srivatsa, Vikranth, et al.
Published: (2024)
Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving
by: Zhao, Zihan, et al.
Published: (2026)
by: Zhao, Zihan, et al.
Published: (2026)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026)
by: Fan, Ruibo, et al.
Published: (2026)
MEPIC: Memory Efficient Position Independent Caching for LLM Serving
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
Progressive Sparse Attention: Algorithm and System Co-design for Efficient Attention in LLM Serving
by: Zhou, Qihui, et al.
Published: (2025)
by: Zhou, Qihui, et al.
Published: (2025)
AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM Serving
by: Xu, Tianhao, et al.
Published: (2026)
by: Xu, Tianhao, et al.
Published: (2026)
From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
by: Zhu, Tianhao, et al.
Published: (2025)
by: Zhu, Tianhao, et al.
Published: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
by: Tao, Yiheng, et al.
Published: (2025)
by: Tao, Yiheng, et al.
Published: (2025)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026)
by: Wu, Yinpeng, et al.
Published: (2026)
Semantic Caching for Low-Cost LLM Serving: From Offline Learning to Online Adaptation
by: Liu, Xutong, et al.
Published: (2025)
by: Liu, Xutong, et al.
Published: (2025)
ATP: Enabling Fast LLM Serving via Attention on Top Principal Keys
by: Niu, Yue, et al.
Published: (2024)
by: Niu, Yue, et al.
Published: (2024)
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
by: Ao, Ruicheng, et al.
Published: (2025)
by: Ao, Ruicheng, et al.
Published: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
by: Lin, Yujun, et al.
Published: (2024)
by: Lin, Yujun, et al.
Published: (2024)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
by: Lou, Chiheng, et al.
Published: (2025)
by: Lou, Chiheng, et al.
Published: (2025)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
by: Bin, Kyungmin, et al.
Published: (2025)
by: Bin, Kyungmin, et al.
Published: (2025)
POLAR: Online Learning for LoRA Adapter Caching and Routing in Edge LLM Serving
by: Li, Shaoang, et al.
Published: (2026)
by: Li, Shaoang, et al.
Published: (2026)
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
by: Gao, Lei, et al.
Published: (2025)
by: Gao, Lei, et al.
Published: (2025)
MRMMIA: Membership Inference Attacks on Memory in Chat Agents
by: Chen, Kai, et al.
Published: (2026)
by: Chen, Kai, et al.
Published: (2026)
Do Not Wait: Learning Re-Ranking Model Without User Feedback At Serving Time in E-Commerce
by: Wang, Yuan, et al.
Published: (2024)
by: Wang, Yuan, et al.
Published: (2024)
FastLogAD: Log Anomaly Detection with Mask-Guided Pseudo Anomaly Generation and Discrimination
by: Lin, Yifei, et al.
Published: (2024)
by: Lin, Yifei, et al.
Published: (2024)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
by: Zhang, Yiqi, et al.
Published: (2026)
by: Zhang, Yiqi, et al.
Published: (2026)
Llumnix: Dynamic Scheduling for Large Language Model Serving
by: Sun, Biao, et al.
Published: (2024)
by: Sun, Biao, et al.
Published: (2024)
Duration Aware Scheduling for ASR Serving Under Workload Drift
by: Makwana, Darshan, et al.
Published: (2026)
by: Makwana, Darshan, et al.
Published: (2026)
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
by: Zhao, Yilong, et al.
Published: (2023)
by: Zhao, Yilong, et al.
Published: (2023)
Towards Fast Safe Online Reinforcement Learning via Policy Finetuning
by: Chen, Keru, et al.
Published: (2024)
by: Chen, Keru, et al.
Published: (2024)
PRISM-CTG: A Foundation Model for Cardiotocography Analysis with Multi-View SSL
by: Wong, Sheng, et al.
Published: (2026)
by: Wong, Sheng, et al.
Published: (2026)
Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
by: Lin, Chaofan, et al.
Published: (2024)
by: Lin, Chaofan, et al.
Published: (2024)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
by: Huang, Shaoyuan, et al.
Published: (2026)
by: Huang, Shaoyuan, et al.
Published: (2026)
FedLGA: Towards System-Heterogeneity of Federated Learning via Local Gradient Approximation
by: Li, Xingyu, et al.
Published: (2021)
by: Li, Xingyu, et al.
Published: (2021)
Data Augmentation for Continual RL via Adversarial Gradient Episodic Memory
by: Wu, Sihao, et al.
Published: (2024)
by: Wu, Sihao, et al.
Published: (2024)
Streaming, Fast and Slow: Cognitive Load-Aware Streaming for Efficient LLM Serving
by: Xiao, Chang, et al.
Published: (2025)
by: Xiao, Chang, et al.
Published: (2025)
Similar Items
-
Echo: Efficient Co-Scheduling of Hybrid Online-Offline Tasks for Large Language Model Serving
by: Wang, Zhibin, et al.
Published: (2025) -
Locality-aware Fair Scheduling in LLM Serving
by: Cao, Shiyi, et al.
Published: (2025) -
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
by: Gao, Shihong, et al.
Published: (2025) -
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
by: Qiao, Yifan, et al.
Published: (2024) -
PRISM: Parallel Residual Iterative Sequence Model
by: Jiang, Jie, et al.
Published: (2026)