ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Zaifeng, Shen, Yipeng, Hu, Zhengding, Wang, Zhuang, Manocha, Aninda, Wang, Zheng, Yu, Zhongkai, Guan, Yue, Ding, Yufei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows
by: Pan, Zaifeng, et al.
Published: (2025)
by: Pan, Zaifeng, et al.
Published: (2025)
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
by: Hu, Zhengding, et al.
Published: (2026)
by: Hu, Zhengding, et al.
Published: (2026)
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
by: Qiang, Xinwei, et al.
Published: (2026)
by: Qiang, Xinwei, et al.
Published: (2026)
AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving
by: Yu, Zhongkai, et al.
Published: (2026)
by: Yu, Zhongkai, et al.
Published: (2026)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
by: Yu, Zhongkai, et al.
Published: (2025)
by: Yu, Zhongkai, et al.
Published: (2025)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
by: Cho, Jaehong, et al.
Published: (2024)
by: Cho, Jaehong, et al.
Published: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
by: Razavi, Kamran, et al.
Published: (2024)
by: Razavi, Kamran, et al.
Published: (2024)
DRackSim: Simulator for Rack-scale Memory Disaggregation
by: Puri, Amit, et al.
Published: (2023)
by: Puri, Amit, et al.
Published: (2023)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
by: Yoon, Dongha, et al.
Published: (2025)
by: Yoon, Dongha, et al.
Published: (2025)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
by: Xu, Jiale, et al.
Published: (2025)
by: Xu, Jiale, et al.
Published: (2025)
Sponge: Inference Serving with Dynamic SLOs Using In-Place Vertical Scaling
by: Razavi, Kamran, et al.
Published: (2024)
by: Razavi, Kamran, et al.
Published: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
by: Shi, Ge, et al.
Published: (2025)
by: Shi, Ge, et al.
Published: (2025)
FLAME: A Serving System Optimized for Large-Scale Generative Recommendation with Efficiency
by: Guo, Xianwen, et al.
Published: (2025)
by: Guo, Xianwen, et al.
Published: (2025)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
by: Nie, Chengyi, et al.
Published: (2024)
by: Nie, Chengyi, et al.
Published: (2024)
KPerfIR: Towards an Open and Compiler-centric Ecosystem for GPU Kernel Performance Tooling on Modern AI Workloads
by: Guan, Yue, et al.
Published: (2025)
by: Guan, Yue, et al.
Published: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
by: Lai, Ruiqi, et al.
Published: (2025)
by: Lai, Ruiqi, et al.
Published: (2025)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
by: Cheng, Rongxin, et al.
Published: (2024)
by: Cheng, Rongxin, et al.
Published: (2024)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
by: Qianli, Liu, et al.
Published: (2025)
by: Qianli, Liu, et al.
Published: (2025)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
by: Jin, Yibo, et al.
Published: (2024)
by: Jin, Yibo, et al.
Published: (2024)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
by: Zhu, Ruidong, et al.
Published: (2025)
by: Zhu, Ruidong, et al.
Published: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
by: Bian, Zhuohang, et al.
Published: (2026)
by: Bian, Zhuohang, et al.
Published: (2026)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
by: Basit, Omar, et al.
Published: (2026)
by: Basit, Omar, et al.
Published: (2026)
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
by: Bu, Tianci, et al.
Published: (2026)
by: Bu, Tianci, et al.
Published: (2026)
PATCHEDSERVE: A Patch Management Framework for SLO-Optimized Hybrid Resolution Diffusion Serving
by: Sun, Desen, et al.
Published: (2025)
by: Sun, Desen, et al.
Published: (2025)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
by: Cho, Jaehong, et al.
Published: (2026)
by: Cho, Jaehong, et al.
Published: (2026)
Exploring the Frontiers of Energy Efficiency using Power Management at System Scale
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
SplitSim: Large-Scale Simulations for Evaluating Network Systems Research
by: Li, Hejing, et al.
Published: (2024)
by: Li, Hejing, et al.
Published: (2024)
Model Input Verification of Large Scale Simulations
by: Neykova, Rumyana, et al.
Published: (2024)
by: Neykova, Rumyana, et al.
Published: (2024)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
by: Wang, Weiye, et al.
Published: (2026)
by: Wang, Weiye, et al.
Published: (2026)
Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
by: Wu, Yebo, et al.
Published: (2024)
by: Wu, Yebo, et al.
Published: (2024)
Similar Items
-
KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows
by: Pan, Zaifeng, et al.
Published: (2025) -
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration
by: Hu, Zhengding, et al.
Published: (2026) -
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
by: Qiang, Xinwei, et al.
Published: (2026) -
AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving
by: Yu, Zhongkai, et al.
Published: (2026) -
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)