Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xu, Minxian, Wu, Jingfeng, Song, Shengye, Srirama, Satish Narayana, Javad, Bahman, Ranjan, Rajiv, Jha, Devki Nandan, Wang, Sa, Tian, Wenhong, Xu, Huanle, Li, Li, Mo, Zizhao, Ren, Shuo, Kunz, Thomas, Kochovski, Petar, Stankovski, Vlado, Ye, Kejiang, Xu, Chengzhong, Buyya, Rajkumar |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
par: Bai, Haoyu, et autres
Publié: (2024)
par: Bai, Haoyu, et autres
Publié: (2024)
A Context‐Aware Decision Support Framework for Scientific Experiment Configuration
par: Pouriya Miri, et autres
Publié: (2026)
par: Pouriya Miri, et autres
Publié: (2026)
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
par: Wu, Jingfeng, et autres
Publié: (2024)
par: Wu, Jingfeng, et autres
Publié: (2024)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
par: Wen, Linfeng, et autres
Publié: (2024)
par: Wen, Linfeng, et autres
Publié: (2024)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
par: Liao, Junhan, et autres
Publié: (2025)
par: Liao, Junhan, et autres
Publié: (2025)
Auto-scaling Approaches for Microservice Applications: A Survey and Taxonomy
par: Xu, Minxian, et autres
Publié: (2025)
par: Xu, Minxian, et autres
Publié: (2025)
Deep Reinforcement Learning (DRL)-based Methods for Serverless Stream Processing Engines: A Vision, Architectural Elements, and Future Directions
par: Read, Maria R., et autres
Publié: (2024)
par: Read, Maria R., et autres
Publié: (2024)
Cloudnativesim: A Toolkit for Modeling and Simulation of Cloud‐Native Applications
par: Jingfeng Wu, et autres
Publié: (2025)
par: Jingfeng Wu, et autres
Publié: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
par: Mo, Zizhao, et autres
Publié: (2026)
par: Mo, Zizhao, et autres
Publié: (2026)
Distributed Edge Analytics in Edge-Fog-Cloud Continuum
par: Srirama, Satish Narayana
Publié: (2024)
par: Srirama, Satish Narayana
Publié: (2024)
Distributed edge analytics in edge‐fog‐cloud continuum
par: Satish Narayana Srirama
Publié: (2024)
par: Satish Narayana Srirama
Publié: (2024)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
par: Zheng, Wanyi, et autres
Publié: (2025)
par: Zheng, Wanyi, et autres
Publié: (2025)
Cloud Native System for LLM Inference Serving
par: Xu, Minxian, et autres
Publié: (2025)
par: Xu, Minxian, et autres
Publié: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
par: He, Yiyuan, et autres
Publié: (2024)
par: He, Yiyuan, et autres
Publié: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
par: Wu, Jingfeng, et autres
Publié: (2025)
par: Wu, Jingfeng, et autres
Publié: (2025)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
par: Kan Hu, et autres
Publié: (2024)
par: Kan Hu, et autres
Publié: (2024)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
par: Tang, Lujie, et autres
Publié: (2024)
par: Tang, Lujie, et autres
Publié: (2024)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
par: Hu, Kan, et autres
Publié: (2024)
par: Hu, Kan, et autres
Publié: (2024)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
par: Hu, Jianmin, et autres
Publié: (2025)
par: Hu, Jianmin, et autres
Publié: (2025)
Quantum autopoiesis: A quantum-informational model of recursive identity and knowledge discovery
par: Stankovski, Vlado
Publié: (2026)
par: Stankovski, Vlado
Publié: (2026)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
par: Mo, Zizhao, et autres
Publié: (2025)
par: Mo, Zizhao, et autres
Publié: (2025)
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
par: Song, Shengye, et autres
Publié: (2025)
par: Song, Shengye, et autres
Publié: (2025)
C‐Koordinator: Interference‐Aware Management for Large‐Scale and Co‐Located Microservice Clusters
par: Shengye Song, et autres
Publié: (2026)
par: Shengye Song, et autres
Publié: (2026)
TD3-Sched: Learning to Orchestrate Container-based Cloud-Edge Resources via Distributed Reinforcement Learning
par: Song, Shengye, et autres
Publié: (2025)
par: Song, Shengye, et autres
Publié: (2025)
Fog enabled distributed training architecture for federated learning
par: Kumar, Aditya, et autres
Publié: (2024)
par: Kumar, Aditya, et autres
Publié: (2024)
FogDEFTKube: Standards‐compliant dynamic deployment of fog service containers
par: Rajesh Thalla, et autres
Publié: (2024)
par: Rajesh Thalla, et autres
Publié: (2024)
ORACL: Optimized Reasoning for Autoscaling via Chain of Thought with LLMs for Microservices
par: Bai, Haoyu, et autres
Publié: (2026)
par: Bai, Haoyu, et autres
Publié: (2026)
Quality of Experience Based Dynamic Path Serverless Data Pipelines in Edge/Fog Computing
par: Sreenivasu Mirampalli, et autres
Publié: (2025)
par: Sreenivasu Mirampalli, et autres
Publié: (2025)
Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
par: Mo, Zizhao, et autres
Publié: (2024)
par: Mo, Zizhao, et autres
Publié: (2024)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
par: Sun, Yifan, et autres
Publié: (2026)
par: Sun, Yifan, et autres
Publié: (2026)
DeF-DReL: Systematic Deployment of Serverless Functions in Fog and Cloud environments using Deep Reinforcement Learning
par: Dehury, Chinmaya Kumar, et autres
Publié: (2021)
par: Dehury, Chinmaya Kumar, et autres
Publié: (2021)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: He, Yiyuan, et autres
Publié: (2025)
par: He, Yiyuan, et autres
Publié: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: Yiyuan He, et autres
Publié: (2026)
par: Yiyuan He, et autres
Publié: (2026)
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
par: Li, Xiang, et autres
Publié: (2024)
par: Li, Xiang, et autres
Publié: (2024)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
par: Hu, Kan, et autres
Publié: (2024)
par: Hu, Kan, et autres
Publié: (2024)
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
par: Wen, Linfeng, et autres
Publié: (2024)
par: Wen, Linfeng, et autres
Publié: (2024)
Peptide Stapling Using Sonogashira Coupling
par: Devki Nandan, et autres
Publié: (2025)
par: Devki Nandan, et autres
Publié: (2025)
Benchmarking of CPU-intensive Stream Data Processing in The Edge Computing Systems
par: Szydlo, Tomasz, et autres
Publié: (2025)
par: Szydlo, Tomasz, et autres
Publié: (2025)
Deep Reinforcement Learning-based Methods for Resource Scheduling in Cloud Computing: A Review and Future Directions
par: Zhou, Guangyao, et autres
Publié: (2021)
par: Zhou, Guangyao, et autres
Publié: (2021)
SealOS+: A Sealos-based Approach for Adaptive Resource Optimization Under Dynamic Workloads for Securities Trading System
par: Jia, Haojie, et autres
Publié: (2025)
par: Jia, Haojie, et autres
Publié: (2025)
Documents similaires
-
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
par: Bai, Haoyu, et autres
Publié: (2024) -
A Context‐Aware Decision Support Framework for Scientific Experiment Configuration
par: Pouriya Miri, et autres
Publié: (2026) -
CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications
par: Wu, Jingfeng, et autres
Publié: (2024) -
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
par: Wen, Linfeng, et autres
Publié: (2024) -
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
par: Liao, Junhan, et autres
Publié: (2025)