Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Jingfeng, He, Yiyuan, Xu, Minxian, Gao, Xitong, Ye, Kejiang, Xu, Chengzhong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Cloud Native System for LLM Inference Serving
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
di: He, Yiyuan, et al.
Pubblicazione: (2024)
di: He, Yiyuan, et al.
Pubblicazione: (2024)
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
di: Wu, Jingfeng, et al.
Pubblicazione: (2024)
di: Wu, Jingfeng, et al.
Pubblicazione: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
di: Liao, Junhan, et al.
Pubblicazione: (2025)
di: Liao, Junhan, et al.
Pubblicazione: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
di: Tang, Lujie, et al.
Pubblicazione: (2024)
di: Tang, Lujie, et al.
Pubblicazione: (2024)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
di: Hu, Kan, et al.
Pubblicazione: (2024)
di: Hu, Kan, et al.
Pubblicazione: (2024)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
di: Bai, Haoyu, et al.
Pubblicazione: (2024)
di: Bai, Haoyu, et al.
Pubblicazione: (2024)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
di: Lin, Yanying, et al.
Pubblicazione: (2025)
di: Lin, Yanying, et al.
Pubblicazione: (2025)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
di: Mo, Zizhao, et al.
Pubblicazione: (2025)
di: Mo, Zizhao, et al.
Pubblicazione: (2025)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
di: Hu, Kan, et al.
Pubblicazione: (2024)
di: Hu, Kan, et al.
Pubblicazione: (2024)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
di: Song, Shengye, et al.
Pubblicazione: (2025)
di: Song, Shengye, et al.
Pubblicazione: (2025)
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
di: Jia, Haojie, et al.
Pubblicazione: (2025)
di: Jia, Haojie, et al.
Pubblicazione: (2025)
TD3-Sched: Learning to Orchestrate Container-based Cloud-Edge Resources via Distributed Reinforcement Learning
di: Song, Shengye, et al.
Pubblicazione: (2025)
di: Song, Shengye, et al.
Pubblicazione: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
di: Li, Cong, et al.
Pubblicazione: (2025)
di: Li, Cong, et al.
Pubblicazione: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
di: Zhou, Qihui, et al.
Pubblicazione: (2025)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
di: Xu, Minxian, et al.
Pubblicazione: (2026)
di: Xu, Minxian, et al.
Pubblicazione: (2026)
BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
di: Hu, Bodun, et al.
Pubblicazione: (2024)
di: Hu, Bodun, et al.
Pubblicazione: (2024)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
di: Chen, Wenyan, et al.
Pubblicazione: (2026)
di: Chen, Wenyan, et al.
Pubblicazione: (2026)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
di: Sun, Yifan, et al.
Pubblicazione: (2026)
di: Sun, Yifan, et al.
Pubblicazione: (2026)
ORACL: Optimized Reasoning for Autoscaling via Chain of Thought with LLMs for Microservices
di: Bai, Haoyu, et al.
Pubblicazione: (2026)
di: Bai, Haoyu, et al.
Pubblicazione: (2026)
DeepServe: Serverless Large Language Model Serving at Scale
di: Hu, Junhao, et al.
Pubblicazione: (2025)
di: Hu, Junhao, et al.
Pubblicazione: (2025)
TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
di: Bian, Zhuohang, et al.
Pubblicazione: (2026)
Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
di: Wu, Yebo, et al.
Pubblicazione: (2024)
di: Wu, Yebo, et al.
Pubblicazione: (2024)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
di: Peng, You, et al.
Pubblicazione: (2026)
di: Peng, You, et al.
Pubblicazione: (2026)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
Fine-grained MoE Load Balancing with Linear Programming
di: Zhao, Chenqi, et al.
Pubblicazione: (2025)
di: Zhao, Chenqi, et al.
Pubblicazione: (2025)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
di: Jiang, Youhe, et al.
Pubblicazione: (2026)
Fast State Restoration in LLM Serving with HCache
di: Gao, Shiwei, et al.
Pubblicazione: (2024)
di: Gao, Shiwei, et al.
Pubblicazione: (2024)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
di: Qiao, Yifan, et al.
Pubblicazione: (2024)
di: Qiao, Yifan, et al.
Pubblicazione: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Cloud Native System for LLM Inference Serving
di: Xu, Minxian, et al.
Pubblicazione: (2025) -
UELLM: A Unified and Efficient Approach for LLM Inference Serving
di: He, Yiyuan, et al.
Pubblicazione: (2024) -
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
di: Wu, Jingfeng, et al.
Pubblicazione: (2024) -
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025) -
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)