LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
Fuente:
arXiv
Salvato in:
| Autori principali: | Cho, Jaehong, Choi, Hyunmin, Heo, Guseul, Park, Jongse |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LLMServingSim2.0: A Unified Simulator for Heterogeneous Hardware and Serving Techniques in LLM Infrastructure
di: Cho, Jaehong, et al.
Pubblicazione: (2025)
di: Cho, Jaehong, et al.
Pubblicazione: (2025)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
di: Cho, Jaehong, et al.
Pubblicazione: (2024)
di: Cho, Jaehong, et al.
Pubblicazione: (2024)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
di: Li, Rongzhi, et al.
Pubblicazione: (2025)
di: Li, Rongzhi, et al.
Pubblicazione: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
di: Qiu, Haoran, et al.
Pubblicazione: (2025)
di: Qiu, Haoran, et al.
Pubblicazione: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
di: Bai, Fan, et al.
Pubblicazione: (2026)
di: Bai, Fan, et al.
Pubblicazione: (2026)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management
di: Pan, Zaifeng, et al.
Pubblicazione: (2026)
di: Pan, Zaifeng, et al.
Pubblicazione: (2026)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
di: Guo, Yipin, et al.
Pubblicazione: (2026)
di: Guo, Yipin, et al.
Pubblicazione: (2026)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
di: Kamahori, Keisuke, et al.
Pubblicazione: (2026)
di: Kamahori, Keisuke, et al.
Pubblicazione: (2026)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
di: Liu, Zedong, et al.
Pubblicazione: (2026)
di: Liu, Zedong, et al.
Pubblicazione: (2026)
DRackSim: Simulator for Rack-scale Memory Disaggregation
di: Puri, Amit, et al.
Pubblicazione: (2023)
di: Puri, Amit, et al.
Pubblicazione: (2023)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
di: Gao, Wei, et al.
Pubblicazione: (2025)
di: Gao, Wei, et al.
Pubblicazione: (2025)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
di: Cheng, Rongxin, et al.
Pubblicazione: (2024)
di: Cheng, Rongxin, et al.
Pubblicazione: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
di: Lee, Sanghyeon, et al.
Pubblicazione: (2025)
di: Lee, Sanghyeon, et al.
Pubblicazione: (2025)
LLM Inference Serving: Survey of Recent Advances and Opportunities
di: Li, Baolin, et al.
Pubblicazione: (2024)
di: Li, Baolin, et al.
Pubblicazione: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
LAPS: A Length-Aware-Prefill LLM Serving System
di: She, Jianshu, et al.
Pubblicazione: (2026)
di: She, Jianshu, et al.
Pubblicazione: (2026)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
di: Wagenländer, Marcel, et al.
Pubblicazione: (2026)
di: Wagenländer, Marcel, et al.
Pubblicazione: (2026)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
di: Hankendi, Can, et al.
Pubblicazione: (2026)
di: Hankendi, Can, et al.
Pubblicazione: (2026)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
di: Zhao, Juntao, et al.
Pubblicazione: (2024)
di: Zhao, Juntao, et al.
Pubblicazione: (2024)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
di: Mitra, Tiyasa, et al.
Pubblicazione: (2025)
di: Mitra, Tiyasa, et al.
Pubblicazione: (2025)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
di: Li, Jiaxi, et al.
Pubblicazione: (2025)
di: Li, Jiaxi, et al.
Pubblicazione: (2025)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
di: Huang, Tao, et al.
Pubblicazione: (2024)
di: Huang, Tao, et al.
Pubblicazione: (2024)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
di: Da, Wei, et al.
Pubblicazione: (2025)
di: Da, Wei, et al.
Pubblicazione: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
di: Wu, Linyu, et al.
Pubblicazione: (2025)
di: Wu, Linyu, et al.
Pubblicazione: (2025)
MinT: Managed Infrastructure for Training and Serving Millions of LLMs
di: Lab, Mind, et al.
Pubblicazione: (2026)
di: Lab, Mind, et al.
Pubblicazione: (2026)
Accelerating LLM Inference with Precomputed Query Storage
di: Park, Jay H., et al.
Pubblicazione: (2025)
di: Park, Jay H., et al.
Pubblicazione: (2025)
Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics
di: Zhou, Zijie
Pubblicazione: (2026)
di: Zhou, Zijie
Pubblicazione: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
di: Pan, Xinglin, et al.
Pubblicazione: (2025)
di: Pan, Xinglin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LLMServingSim2.0: A Unified Simulator for Heterogeneous Hardware and Serving Techniques in LLM Infrastructure
di: Cho, Jaehong, et al.
Pubblicazione: (2025) -
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
di: Cho, Jaehong, et al.
Pubblicazione: (2024) -
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
di: Kumar, Satyam, et al.
Pubblicazione: (2026) -
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
di: Li, Rongzhi, et al.
Pubblicazione: (2025) -
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)