Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Chendong, Wang, Meixuan, Zhou, Hang, Liang, Hong, Lyu, Yuan, Chen, Zixi, Fan, Yuwei, Zhou, Zijie |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
by: Bu, Tianci, et al.
Published: (2026)
by: Bu, Tianci, et al.
Published: (2026)
LLM Serving Optimization with Variable Prefill and Decode Lengths
by: Wang, Meixuan, et al.
Published: (2025)
by: Wang, Meixuan, et al.
Published: (2025)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
by: Wu, Hanjiang, et al.
Published: (2026)
by: Wu, Hanjiang, et al.
Published: (2026)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025)
by: Chen, Zixi, et al.
Published: (2025)
TreeGPT: Pure TreeFFN Encoder-Decoder Architecture for Structured Reasoning Without Attention Mechanisms
by: Li, Zixi
Published: (2025)
by: Li, Zixi
Published: (2025)
A Universal Load Balancing Principle and Its Application to Large Language Model Serving
by: Chen, Zixi, et al.
Published: (2026)
by: Chen, Zixi, et al.
Published: (2026)
Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics
by: Zhou, Zijie
Published: (2026)
by: Zhou, Zijie
Published: (2026)
Revealing the Challenges of Attention-FFN Disaggregation for Modern MoE Models and Hardware Systems
by: Liu, Guowei, et al.
Published: (2026)
by: Liu, Guowei, et al.
Published: (2026)
UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
by: Zhou, Hanzhang, et al.
Published: (2024)
by: Zhou, Hanzhang, et al.
Published: (2024)
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
by: Liang, Yunkai, et al.
Published: (2025)
by: Liang, Yunkai, et al.
Published: (2025)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
by: Kumar, Satyam, et al.
Published: (2026)
by: Kumar, Satyam, et al.
Published: (2026)
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
by: Pei, Zehua, et al.
Published: (2025)
by: Pei, Zehua, et al.
Published: (2025)
UMoE: Unifying Attention and FFN with Shared Experts
by: Yang, Yuanhang, et al.
Published: (2025)
by: Yang, Yuanhang, et al.
Published: (2025)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
by: Qin, Ruoyu, et al.
Published: (2024)
by: Qin, Ruoyu, et al.
Published: (2024)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
by: Bai, Fan, et al.
Published: (2026)
by: Bai, Fan, et al.
Published: (2026)
Sparse-VQ Transformer: An FFN-Free Framework with Vector Quantization for Enhanced Time Series Forecasting
by: Zhao, Yanjun, et al.
Published: (2024)
by: Zhao, Yanjun, et al.
Published: (2024)
Progressive Sparse Attention: Algorithm and System Co-design for Efficient Attention in LLM Serving
by: Zhou, Qihui, et al.
Published: (2025)
by: Zhou, Qihui, et al.
Published: (2025)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
by: Guo, Yipin, et al.
Published: (2026)
by: Guo, Yipin, et al.
Published: (2026)
Fast Forward: Accelerating LLM Prefill with Predictive FFN Sparsity
by: Gautam, Aayush, et al.
Published: (2026)
by: Gautam, Aayush, et al.
Published: (2026)
On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
by: Huang, Jen-tse, et al.
Published: (2024)
by: Huang, Jen-tse, et al.
Published: (2024)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
by: Liu, Zedong, et al.
Published: (2026)
by: Liu, Zedong, et al.
Published: (2026)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
by: Cho, Jaehong, et al.
Published: (2026)
by: Cho, Jaehong, et al.
Published: (2026)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
by: Qiu, Haoran, et al.
Published: (2025)
by: Qiu, Haoran, et al.
Published: (2025)
Efficiently Serving Large Multimodal Models Using EPD Disaggregation
by: Singh, Gursimran, et al.
Published: (2024)
by: Singh, Gursimran, et al.
Published: (2024)
Sparsity Moves Computation: How FFN Architecture Reshapes Attention in Small Transformers
by: Smithline, Gabriel, et al.
Published: (2026)
by: Smithline, Gabriel, et al.
Published: (2026)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
by: Wang, Haodong, et al.
Published: (2025)
by: Wang, Haodong, et al.
Published: (2025)
Trinity: Disaggregating Vector Search from Prefill-Decode Disaggregation in LLM Serving
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
by: Ruan, Chaoyi, et al.
Published: (2025)
by: Ruan, Chaoyi, et al.
Published: (2025)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
by: Hu, Xiannan, et al.
Published: (2025)
by: Hu, Xiannan, et al.
Published: (2025)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
PDTrim: Targeted Pruning for Prefill-Decode Disaggregation in Inference
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
by: Ye, Zihao, et al.
Published: (2025)
by: Ye, Zihao, et al.
Published: (2025)
SOMA: Efficient Multi-turn LLM Serving via Small Language Model
by: Cheng, Xueqi, et al.
Published: (2026)
by: Cheng, Xueqi, et al.
Published: (2026)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
by: Yang, Shang, et al.
Published: (2025)
by: Yang, Shang, et al.
Published: (2025)
QMGeo: Differentially Private Federated Learning via Stochastic Quantization with Mixed Truncated Geometric Distribution
by: Wang, Zixi, et al.
Published: (2023)
by: Wang, Zixi, et al.
Published: (2023)
Similar Items
-
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
by: Bu, Tianci, et al.
Published: (2026) -
LLM Serving Optimization with Variable Prefill and Decode Lengths
by: Wang, Meixuan, et al.
Published: (2025) -
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
by: Wu, Hanjiang, et al.
Published: (2026) -
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025) -
TreeGPT: Pure TreeFFN Encoder-Decoder Architecture for Structured Reasoning Without Attention Mechanisms
by: Li, Zixi
Published: (2025)