LLMServingSim2.0: A Unified Simulator for Heterogeneous Hardware and Serving Techniques in LLM Infrastructure
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cho, Jaehong, Choi, Hyunmin, Park, Jongse |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
von: Cho, Jaehong, et al.
Veröffentlicht: (2026)
von: Cho, Jaehong, et al.
Veröffentlicht: (2026)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
von: Cho, Jaehong, et al.
Veröffentlicht: (2024)
von: Cho, Jaehong, et al.
Veröffentlicht: (2024)
ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management
von: Pan, Zaifeng, et al.
Veröffentlicht: (2026)
von: Pan, Zaifeng, et al.
Veröffentlicht: (2026)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
LAPS: A Length-Aware-Prefill LLM Serving System
von: She, Jianshu, et al.
Veröffentlicht: (2026)
von: She, Jianshu, et al.
Veröffentlicht: (2026)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
von: Huang, Tao, et al.
Veröffentlicht: (2024)
von: Huang, Tao, et al.
Veröffentlicht: (2024)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
von: Da, Wei, et al.
Veröffentlicht: (2025)
von: Da, Wei, et al.
Veröffentlicht: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
MinT: Managed Infrastructure for Training and Serving Millions of LLMs
von: Lab, Mind, et al.
Veröffentlicht: (2026)
von: Lab, Mind, et al.
Veröffentlicht: (2026)
Accelerating LLM Inference with Precomputed Query Storage
von: Park, Jay H., et al.
Veröffentlicht: (2025)
von: Park, Jay H., et al.
Veröffentlicht: (2025)
Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics
von: Zhou, Zijie
Veröffentlicht: (2026)
von: Zhou, Zijie
Veröffentlicht: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
von: Jayakody, Shakya, et al.
Veröffentlicht: (2026)
von: Jayakody, Shakya, et al.
Veröffentlicht: (2026)
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
von: Hu, Huanqi, et al.
Veröffentlicht: (2025)
von: Hu, Huanqi, et al.
Veröffentlicht: (2025)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
von: Song, Xiaoniu, et al.
Veröffentlicht: (2024)
von: Song, Xiaoniu, et al.
Veröffentlicht: (2024)
High-Throughput LLM inference on Heterogeneous Clusters
von: Xiong, Yi, et al.
Veröffentlicht: (2025)
von: Xiong, Yi, et al.
Veröffentlicht: (2025)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
von: Mao, Ziming, et al.
Veröffentlicht: (2024)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving
von: Hsieh, Chia-chi, et al.
Veröffentlicht: (2026)
von: Hsieh, Chia-chi, et al.
Veröffentlicht: (2026)
Trillion Parameter AI Serving Infrastructure for Scientific Discovery: A Survey and Vision
von: Hudson, Nathaniel, et al.
Veröffentlicht: (2024)
von: Hudson, Nathaniel, et al.
Veröffentlicht: (2024)
PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving
von: Fang, Fei, et al.
Veröffentlicht: (2025)
von: Fang, Fei, et al.
Veröffentlicht: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
Regulating Branch Parallelism in LLM Serving
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
Byzantine-Robust Decentralized Coordination of LLM Agents
von: Jo, Yongrae, et al.
Veröffentlicht: (2025)
von: Jo, Yongrae, et al.
Veröffentlicht: (2025)
Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models
von: Chen, Daoyuan, et al.
Veröffentlicht: (2024)
von: Chen, Daoyuan, et al.
Veröffentlicht: (2024)
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures
von: Kilic, Ozgur O., et al.
Veröffentlicht: (2025)
von: Kilic, Ozgur O., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
von: Cho, Jaehong, et al.
Veröffentlicht: (2026) -
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
von: Cho, Jaehong, et al.
Veröffentlicht: (2024) -
ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management
von: Pan, Zaifeng, et al.
Veröffentlicht: (2026) -
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
von: Yuan, Yichao, et al.
Veröffentlicht: (2025) -
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)