DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
Fuente:
arXiv
Guardado en:
| Autores principales: | Gao, Lei, Jiang, Chaoyi, Zarch, Hossein Entezari, Wong, Daniel, Hill, Mark, Annavaram, Murali |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
por: Zarch, Hossein Entezari, et al.
Publicado: (2025)
por: Zarch, Hossein Entezari, et al.
Publicado: (2025)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
por: Jiang, Chaoyi, et al.
Publicado: (2024)
por: Jiang, Chaoyi, et al.
Publicado: (2024)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
por: Zarch, Hossein Entezari, et al.
Publicado: (2025)
por: Zarch, Hossein Entezari, et al.
Publicado: (2025)
CADC: Encoding User-Item Interactions for Compressing Recommendation Model Training Data
por: Zarch, Hossein Entezari, et al.
Publicado: (2024)
por: Zarch, Hossein Entezari, et al.
Publicado: (2024)
MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
por: Jiang, Chaoyi, et al.
Publicado: (2025)
por: Jiang, Chaoyi, et al.
Publicado: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
por: Ramachandran, Arun, et al.
Publicado: (2025)
por: Ramachandran, Arun, et al.
Publicado: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
por: Shi, Xiaoxiang, et al.
Publicado: (2025)
por: Shi, Xiaoxiang, et al.
Publicado: (2025)
LLM Serving Optimization with Variable Prefill and Decode Lengths
por: Wang, Meixuan, et al.
Publicado: (2025)
por: Wang, Meixuan, et al.
Publicado: (2025)
Fast NF4 Dequantization Kernels for Large Language Model Inference
por: Qi, Xiangbo, et al.
Publicado: (2026)
por: Qi, Xiangbo, et al.
Publicado: (2026)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
por: Chen, Yukang, et al.
Publicado: (2025)
por: Chen, Yukang, et al.
Publicado: (2025)
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
por: Gao, Shihong, et al.
Publicado: (2025)
por: Gao, Shihong, et al.
Publicado: (2025)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
por: Woo, Sunghyeon, et al.
Publicado: (2026)
por: Woo, Sunghyeon, et al.
Publicado: (2026)
PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior
por: Flemings, James, et al.
Publicado: (2026)
por: Flemings, James, et al.
Publicado: (2026)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
por: Lou, Chiheng, et al.
Publicado: (2025)
por: Lou, Chiheng, et al.
Publicado: (2025)
Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels
por: Song, Mingcong, et al.
Publicado: (2024)
por: Song, Mingcong, et al.
Publicado: (2024)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
por: Kong, Linghao, et al.
Publicado: (2026)
por: Kong, Linghao, et al.
Publicado: (2026)
MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving
por: Su, Zhaoyuan, et al.
Publicado: (2026)
por: Su, Zhaoyuan, et al.
Publicado: (2026)
MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines
por: Gao, Lei, et al.
Publicado: (2024)
por: Gao, Lei, et al.
Publicado: (2024)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
por: Qiao, Yifan, et al.
Publicado: (2024)
por: Qiao, Yifan, et al.
Publicado: (2024)
Differentially Private Knowledge Distillation via Synthetic Text Generation
por: Flemings, James, et al.
Publicado: (2024)
por: Flemings, James, et al.
Publicado: (2024)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
por: Zhong, Yinmin, et al.
Publicado: (2024)
por: Zhong, Yinmin, et al.
Publicado: (2024)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
por: Wang, Chao, et al.
Publicado: (2025)
por: Wang, Chao, et al.
Publicado: (2025)
Trinity: Disaggregating Vector Search from Prefill-Decode Disaggregation in LLM Serving
por: Liu, Yi, et al.
Publicado: (2025)
por: Liu, Yi, et al.
Publicado: (2025)
Adaptively Private Next-Token Prediction of Large Language Models
por: Flemings, James, et al.
Publicado: (2024)
por: Flemings, James, et al.
Publicado: (2024)
DistilLock: Safeguarding LLMs from Unauthorized Knowledge Distillation on the Edge
por: Mohanty, Asmita, et al.
Publicado: (2025)
por: Mohanty, Asmita, et al.
Publicado: (2025)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
por: Li, Zikun, et al.
Publicado: (2025)
por: Li, Zikun, et al.
Publicado: (2025)
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
por: Agrawal, Amey, et al.
Publicado: (2026)
por: Agrawal, Amey, et al.
Publicado: (2026)
LAPS: A Length-Aware-Prefill LLM Serving System
por: She, Jianshu, et al.
Publicado: (2026)
por: She, Jianshu, et al.
Publicado: (2026)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
por: Yu, Shan, et al.
Publicado: (2025)
por: Yu, Shan, et al.
Publicado: (2025)
Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
por: Agullo, Ferran, et al.
Publicado: (2026)
por: Agullo, Ferran, et al.
Publicado: (2026)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
por: Duan, Jiangfei, et al.
Publicado: (2024)
por: Duan, Jiangfei, et al.
Publicado: (2024)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
por: Lee, Gunjun, et al.
Publicado: (2025)
por: Lee, Gunjun, et al.
Publicado: (2025)
Efficient Mixture-of-Agents Serving via Tree-Structured Routing, Adaptive Pruning, and Dependency-Aware Prefill-Decode Overlap
por: Wang, Zijun, et al.
Publicado: (2025)
por: Wang, Zijun, et al.
Publicado: (2025)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
por: Huang, Shaoyuan, et al.
Publicado: (2026)
por: Huang, Shaoyuan, et al.
Publicado: (2026)
Continuous Semantic Caching for Low-Cost LLM Serving
por: Atalar, Baran, et al.
Publicado: (2026)
por: Atalar, Baran, et al.
Publicado: (2026)
Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
por: Li, Zongze, et al.
Publicado: (2026)
por: Li, Zongze, et al.
Publicado: (2026)
MPC-Pipe: an Efficient Pipeline Scheme for Secure Multi-party Machine Learning Inference
por: Wang, Yongqin, et al.
Publicado: (2022)
por: Wang, Yongqin, et al.
Publicado: (2022)
LRD-MPC: Efficient MPC Inference through Low-rank Decomposition
por: Tang, Tingting, et al.
Publicado: (2026)
por: Tang, Tingting, et al.
Publicado: (2026)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
por: Kakolyris, Andreas Kosmas, et al.
Publicado: (2024)
por: Kakolyris, Andreas Kosmas, et al.
Publicado: (2024)
CascadeServe: Unlocking Model Cascades for Inference Serving
por: Kossmann, Ferdi, et al.
Publicado: (2024)
por: Kossmann, Ferdi, et al.
Publicado: (2024)
Ejemplares similares
-
DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding
por: Zarch, Hossein Entezari, et al.
Publicado: (2025) -
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
por: Jiang, Chaoyi, et al.
Publicado: (2024) -
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
por: Zarch, Hossein Entezari, et al.
Publicado: (2025) -
CADC: Encoding User-Item Interactions for Compressing Recommendation Model Training Data
por: Zarch, Hossein Entezari, et al.
Publicado: (2024) -
MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
por: Jiang, Chaoyi, et al.
Publicado: (2025)