PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
Fuente:
arXiv
Salvato in:
| Autori principali: | Woo, Sunghyeon, Kim, Hoseung, Shim, Sunghwan, Jo, Minjung, Jeong, Hyunjoon, Lee, Jeongtae, Kim, Joonghoon, Lee, Sungjae, Park, Baeseong, Kwon, Se Jung, Lee, Dongsoo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
di: Lee, Jeongtae, et al.
Pubblicazione: (2026)
di: Lee, Jeongtae, et al.
Pubblicazione: (2026)
ICaRus: Identical Cache Reuse for Efficient Multi Model Inference
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026)
DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward Propagation
di: Woo, Sunghyeon, et al.
Pubblicazione: (2024)
di: Woo, Sunghyeon, et al.
Pubblicazione: (2024)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
di: Park, Gunho, et al.
Pubblicazione: (2022)
di: Park, Gunho, et al.
Pubblicazione: (2022)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
Trinity: Disaggregating Vector Search from Prefill-Decode Disaggregation in LLM Serving
di: Liu, Yi, et al.
Pubblicazione: (2025)
di: Liu, Yi, et al.
Pubblicazione: (2025)
Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
di: Bae, Jeongin, et al.
Pubblicazione: (2026)
di: Bae, Jeongin, et al.
Pubblicazione: (2026)
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection
di: Song, Jiwon, et al.
Pubblicazione: (2026)
di: Song, Jiwon, et al.
Pubblicazione: (2026)
Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
di: Li, Zongze, et al.
Pubblicazione: (2026)
di: Li, Zongze, et al.
Pubblicazione: (2026)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
di: Shi, Xiaoxiang, et al.
Pubblicazione: (2025)
di: Shi, Xiaoxiang, et al.
Pubblicazione: (2025)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
di: Lee, Gunjun, et al.
Pubblicazione: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
QUOKA: Query-Oriented KV Selection For Efficient LLM Prefill
di: Jones, Dalton, et al.
Pubblicazione: (2026)
di: Jones, Dalton, et al.
Pubblicazione: (2026)
MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving
di: Su, Zhaoyuan, et al.
Pubblicazione: (2026)
di: Su, Zhaoyuan, et al.
Pubblicazione: (2026)
DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages
di: Kanani, Alish, et al.
Pubblicazione: (2026)
di: Kanani, Alish, et al.
Pubblicazione: (2026)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
di: Zhang, Hao, et al.
Pubblicazione: (2025)
di: Zhang, Hao, et al.
Pubblicazione: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
di: Chen, Xing, et al.
Pubblicazione: (2025)
di: Chen, Xing, et al.
Pubblicazione: (2025)
Low-Latency Edge LLM Handover via Joint KV Cache Transfer and Token Prefill
di: Lee, Seunghun, et al.
Pubblicazione: (2026)
di: Lee, Seunghun, et al.
Pubblicazione: (2026)
SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
di: Zhang, Hengrui, et al.
Pubblicazione: (2025)
di: Zhang, Hengrui, et al.
Pubblicazione: (2025)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
Accelerating Prefilling for Long-Context LLMs via Sparse Pattern Sharing
di: Peng, Dan, et al.
Pubblicazione: (2025)
di: Peng, Dan, et al.
Pubblicazione: (2025)
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
di: Hsieh, Chia-chi, et al.
Pubblicazione: (2026)
di: Hsieh, Chia-chi, et al.
Pubblicazione: (2026)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
di: Lee, Jung Hyun, et al.
Pubblicazione: (2023)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2023)
LLM Serving Optimization with Variable Prefill and Decode Lengths
di: Wang, Meixuan, et al.
Pubblicazione: (2025)
di: Wang, Meixuan, et al.
Pubblicazione: (2025)
Shallow Prefill, Deep Decoding: Efficient Long-Context Inference via Layer-Asymmetric KV Visibility
di: Oh, Jungsuk, et al.
Pubblicazione: (2026)
di: Oh, Jungsuk, et al.
Pubblicazione: (2026)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
di: Zou, Jing, et al.
Pubblicazione: (2026)
di: Zou, Jing, et al.
Pubblicazione: (2026)
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
di: Tang, Xiaojuan, et al.
Pubblicazione: (2025)
di: Tang, Xiaojuan, et al.
Pubblicazione: (2025)
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
di: Li, Luchang, et al.
Pubblicazione: (2026)
di: Li, Luchang, et al.
Pubblicazione: (2026)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
di: Lee, Joonhyung, et al.
Pubblicazione: (2024)
LAPS: A Length-Aware-Prefill LLM Serving System
di: She, Jianshu, et al.
Pubblicazione: (2026)
di: She, Jianshu, et al.
Pubblicazione: (2026)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
di: Chen, Yukang, et al.
Pubblicazione: (2025)
di: Chen, Yukang, et al.
Pubblicazione: (2025)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
di: Liu, Yunzhao, et al.
Pubblicazione: (2025)
di: Liu, Yunzhao, et al.
Pubblicazione: (2025)
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
di: Gao, Lei, et al.
Pubblicazione: (2025)
di: Gao, Lei, et al.
Pubblicazione: (2025)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
di: Yang, June Yong, et al.
Pubblicazione: (2024)
di: Yang, June Yong, et al.
Pubblicazione: (2024)
FAST-Prefill: FPGA Accelerated Sparse Attention for Long Context LLM Prefill
di: Jayanth, Rakshith, et al.
Pubblicazione: (2026)
di: Jayanth, Rakshith, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026) -
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
di: Lee, Jeongtae, et al.
Pubblicazione: (2026) -
ICaRus: Identical Cache Reuse for Efficient Multi Model Inference
di: Woo, Sunghyeon, et al.
Pubblicazione: (2026) -
DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward Propagation
di: Woo, Sunghyeon, et al.
Pubblicazione: (2024) -
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)