CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Jingfeng, Xu, Minxian, He, Yiyuan, Ye, Kejiang, Xu, Chengzhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
von: Tang, Lujie, et al.
Veröffentlicht: (2024)
von: Tang, Lujie, et al.
Veröffentlicht: (2024)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
von: Hu, Kan, et al.
Veröffentlicht: (2024)
von: Hu, Kan, et al.
Veröffentlicht: (2024)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
von: Bai, Haoyu, et al.
Veröffentlicht: (2024)
von: Bai, Haoyu, et al.
Veröffentlicht: (2024)
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
von: Wen, Linfeng, et al.
Veröffentlicht: (2024)
von: Wen, Linfeng, et al.
Veröffentlicht: (2024)
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
von: Xu, Minxian, et al.
Veröffentlicht: (2026)
von: Xu, Minxian, et al.
Veröffentlicht: (2026)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
von: Liao, Junhan, et al.
Veröffentlicht: (2025)
TD3-Sched: Learning to Orchestrate Container-based Cloud-Edge Resources via Distributed Reinforcement Learning
von: Song, Shengye, et al.
Veröffentlicht: (2025)
von: Song, Shengye, et al.
Veröffentlicht: (2025)
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
von: Hu, Kan, et al.
Veröffentlicht: (2024)
von: Hu, Kan, et al.
Veröffentlicht: (2024)
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
von: Song, Shengye, et al.
Veröffentlicht: (2025)
von: Song, Shengye, et al.
Veröffentlicht: (2025)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
von: Wen, Linfeng, et al.
Veröffentlicht: (2024)
von: Wen, Linfeng, et al.
Veröffentlicht: (2024)
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
von: Jia, Haojie, et al.
Veröffentlicht: (2025)
von: Jia, Haojie, et al.
Veröffentlicht: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
von: Lin, Yanying, et al.
Veröffentlicht: (2025)
Towards cloud-native scientific workflow management
von: Orzechowski, Michal, et al.
Veröffentlicht: (2024)
von: Orzechowski, Michal, et al.
Veröffentlicht: (2024)
Dflow, a Python framework for constructing cloud-native AI-for-Science workflows
von: Liu, Xinzijian, et al.
Veröffentlicht: (2024)
von: Liu, Xinzijian, et al.
Veröffentlicht: (2024)
ORACL: Optimized Reasoning for Autoscaling via Chain of Thought with LLMs for Microservices
von: Bai, Haoyu, et al.
Veröffentlicht: (2026)
von: Bai, Haoyu, et al.
Veröffentlicht: (2026)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
von: Mo, Zizhao, et al.
Veröffentlicht: (2026)
Sarus Suite: Cloud-native Containers for HPC
von: Madonna, Alberto, et al.
Veröffentlicht: (2026)
von: Madonna, Alberto, et al.
Veröffentlicht: (2026)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
von: Sun, Yifan, et al.
Veröffentlicht: (2026)
von: Sun, Yifan, et al.
Veröffentlicht: (2026)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
von: Mo, Zizhao, et al.
Veröffentlicht: (2025)
von: Mo, Zizhao, et al.
Veröffentlicht: (2025)
Edge AI: A Taxonomy, Systematic Review and Future Directions
von: Gill, Sukhpal Singh, et al.
Veröffentlicht: (2024)
von: Gill, Sukhpal Singh, et al.
Veröffentlicht: (2024)
Towards an Adaptive Runtime System for Cloud-Native HPC
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
von: Chazapis, Antony, et al.
Veröffentlicht: (2024)
von: Chazapis, Antony, et al.
Veröffentlicht: (2024)
SimDC: A High-Fidelity Device Simulation Platform for Device-Cloud Collaborative Computing
von: Pei, Ruiguang, et al.
Veröffentlicht: (2025)
von: Pei, Ruiguang, et al.
Veröffentlicht: (2025)
Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
von: Wu, Yebo, et al.
Veröffentlicht: (2024)
von: Wu, Yebo, et al.
Veröffentlicht: (2024)
Service-Level Energy Modeling and Experimentation for Cloud-Native Microservices
von: Legler, Julian, et al.
Veröffentlicht: (2025)
von: Legler, Julian, et al.
Veröffentlicht: (2025)
Object Abstraction To Streamline Edge-Cloud-Native Application Development
von: Lertpongrujikorn, Pawissanutt
Veröffentlicht: (2025)
von: Lertpongrujikorn, Pawissanutt
Veröffentlicht: (2025)
Breaking the Memory Wall for Heterogeneous Federated Learning via Progressive Training
von: Wu, Yebo, et al.
Veröffentlicht: (2024)
von: Wu, Yebo, et al.
Veröffentlicht: (2024)
Artifact for Service-Level Energy Modeling and Experimentation for Cloud-Native Microservices
von: Legler, Julian
Veröffentlicht: (2026)
von: Legler, Julian
Veröffentlicht: (2026)
Optimizing the Longhorn Cloud-native Software Defined Storage Engine for High Performance
von: Kampadais, Konstantinos, et al.
Veröffentlicht: (2025)
von: Kampadais, Konstantinos, et al.
Veröffentlicht: (2025)
On-the-fly Communication-and-Computing to Enable Representation Learning for Distributed Point Clouds
von: Chen, Xu, et al.
Veröffentlicht: (2024)
von: Chen, Xu, et al.
Veröffentlicht: (2024)
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
von: Wu, Feiyang, et al.
Veröffentlicht: (2025)
Computing in the Era of Large Generative Models: From Cloud-Native to AI-Native
von: Lu, Yao, et al.
Veröffentlicht: (2024)
von: Lu, Yao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025) -
UELLM: A Unified and Efficient Approach for LLM Inference Serving
von: He, Yiyuan, et al.
Veröffentlicht: (2024) -
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025) -
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
von: Tang, Lujie, et al.
Veröffentlicht: (2024) -
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
von: Hu, Kan, et al.
Veröffentlicht: (2024)