Efficient Multi-round LLM Inference over Disaggregated Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Wenhao, Jiang, Youhe, Zhao, Penghao, Xu, Quanqing, Yoneki, Eiko, Cui, Bin, Fu, Fangcheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Cascadia: An Efficient Cascade Serving System for Large Language Models
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
von: He, Guoliang, et al.
Veröffentlicht: (2025)
von: He, Guoliang, et al.
Veröffentlicht: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
Parallax: Efficient LLM Inference Service over Decentralized Environment
von: Tong, Chris, et al.
Veröffentlicht: (2025)
von: Tong, Chris, et al.
Veröffentlicht: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
von: Peng, You, et al.
Veröffentlicht: (2026)
von: Peng, You, et al.
Veröffentlicht: (2026)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
von: Ren, Feng, et al.
Veröffentlicht: (2026)
von: Ren, Feng, et al.
Veröffentlicht: (2026)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
von: Basit, Omar, et al.
Veröffentlicht: (2026)
von: Basit, Omar, et al.
Veröffentlicht: (2026)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
von: Chen, Hongyu, et al.
Veröffentlicht: (2026)
LobRA: Multi-tenant Fine-tuning over Heterogeneous Data
von: Lin, Sheng, et al.
Veröffentlicht: (2025)
von: Lin, Sheng, et al.
Veröffentlicht: (2025)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
von: Lai, Ruiqi, et al.
Veröffentlicht: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
von: Yin, Peiqi, et al.
Veröffentlicht: (2026)
von: Yin, Peiqi, et al.
Veröffentlicht: (2026)
KVDirect: Distributed Disaggregated LLM Inference
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
von: Kumar, Madabattula Rajesh, et al.
Veröffentlicht: (2025)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
von: Wu, Yu, et al.
Veröffentlicht: (2025)
von: Wu, Yu, et al.
Veröffentlicht: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
Unleashing Efficient Asynchronous RL Post-Training via Staleness-Constrained Rollout Coordination
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
Improving Automatic Parallel Training via Balanced Memory Workload Optimization
von: Wang, Yujie, et al.
Veröffentlicht: (2023)
von: Wang, Yujie, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
von: Jiang, Youhe, et al.
Veröffentlicht: (2026) -
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
von: Jiang, Youhe, et al.
Veröffentlicht: (2026) -
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
von: Jiang, Youhe, et al.
Veröffentlicht: (2025) -
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025) -
Cascadia: An Efficient Cascade Serving System for Large Language Models
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)