Cloud Native System for LLM Inference Serving
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Minxian, Liao, Junhan, Wu, Jingfeng, He, Yiyuan, Ye, Kejiang, Xu, Chengzhong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
di: Wu, Jingfeng, et al.
Pubblicazione: (2024)
di: Wu, Jingfeng, et al.
Pubblicazione: (2024)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
di: He, Yiyuan, et al.
Pubblicazione: (2024)
di: He, Yiyuan, et al.
Pubblicazione: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
di: Liao, Junhan, et al.
Pubblicazione: (2025)
di: Liao, Junhan, et al.
Pubblicazione: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
di: Tang, Lujie, et al.
Pubblicazione: (2024)
di: Tang, Lujie, et al.
Pubblicazione: (2024)
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
di: Zheng, Wanyi, et al.
Pubblicazione: (2025)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
di: Hu, Kan, et al.
Pubblicazione: (2024)
di: Hu, Kan, et al.
Pubblicazione: (2024)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
di: Bai, Haoyu, et al.
Pubblicazione: (2024)
di: Bai, Haoyu, et al.
Pubblicazione: (2024)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
di: Lin, Yanying, et al.
Pubblicazione: (2025)
di: Lin, Yanying, et al.
Pubblicazione: (2025)
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
TD3-Sched: Learning to Orchestrate Container-based Cloud-Edge Resources via Distributed Reinforcement Learning
di: Song, Shengye, et al.
Pubblicazione: (2025)
di: Song, Shengye, et al.
Pubblicazione: (2025)
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
di: Hu, Kan, et al.
Pubblicazione: (2024)
di: Hu, Kan, et al.
Pubblicazione: (2024)
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
di: Song, Shengye, et al.
Pubblicazione: (2025)
di: Song, Shengye, et al.
Pubblicazione: (2025)
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
di: Jia, Haojie, et al.
Pubblicazione: (2025)
di: Jia, Haojie, et al.
Pubblicazione: (2025)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
di: Wen, Linfeng, et al.
Pubblicazione: (2024)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
di: Xu, Minxian, et al.
Pubblicazione: (2026)
di: Xu, Minxian, et al.
Pubblicazione: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
di: Mo, Zizhao, et al.
Pubblicazione: (2026)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
di: Mo, Zizhao, et al.
Pubblicazione: (2025)
di: Mo, Zizhao, et al.
Pubblicazione: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
di: Zhan, Huiyou, et al.
Pubblicazione: (2025)
di: Zhan, Huiyou, et al.
Pubblicazione: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
di: Gao, Luyao, et al.
Pubblicazione: (2025)
di: Gao, Luyao, et al.
Pubblicazione: (2025)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
di: Jiang, Youhe, et al.
Pubblicazione: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
di: Li, Suyi, et al.
Pubblicazione: (2024)
di: Li, Suyi, et al.
Pubblicazione: (2024)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
di: Lou, Chiheng, et al.
Pubblicazione: (2025)
di: Lou, Chiheng, et al.
Pubblicazione: (2025)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
di: Sun, Yifan, et al.
Pubblicazione: (2026)
di: Sun, Yifan, et al.
Pubblicazione: (2026)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
di: Du, Boxiao, et al.
Pubblicazione: (2026)
di: Du, Boxiao, et al.
Pubblicazione: (2026)
Serving Compound Inference Systems on Datacenter GPUs
di: Devata, Sriram, et al.
Pubblicazione: (2026)
di: Devata, Sriram, et al.
Pubblicazione: (2026)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
di: Zhao, Bohan, et al.
Pubblicazione: (2025)
di: Zhao, Bohan, et al.
Pubblicazione: (2025)
Towards an Adaptive Runtime System for Cloud-Native HPC
di: Bhosale, Aditya, et al.
Pubblicazione: (2026)
di: Bhosale, Aditya, et al.
Pubblicazione: (2026)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
di: Wang, Weiye, et al.
Pubblicazione: (2026)
di: Wang, Weiye, et al.
Pubblicazione: (2026)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
di: Wu, Linyu, et al.
Pubblicazione: (2025)
di: Wu, Linyu, et al.
Pubblicazione: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
di: Ghosh, Himel
Pubblicazione: (2024)
di: Ghosh, Himel
Pubblicazione: (2024)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
di: Da, Wei, et al.
Pubblicazione: (2026)
di: Da, Wei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
di: Wu, Jingfeng, et al.
Pubblicazione: (2024) -
UELLM: A Unified and Efficient Approach for LLM Inference Serving
di: He, Yiyuan, et al.
Pubblicazione: (2024) -
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
di: Wu, Jingfeng, et al.
Pubblicazione: (2025) -
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
di: Liao, Junhan, et al.
Pubblicazione: (2025) -
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)