Decouple and Decompose: Scaling Resource Allocation with DeDe
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Zhiying, Yu, Minlan, Yan, Francis Y. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
von: Hu, Kan, et al.
Veröffentlicht: (2024)
von: Hu, Kan, et al.
Veröffentlicht: (2024)
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
von: Huang, You-Liang, et al.
Veröffentlicht: (2026)
von: Huang, You-Liang, et al.
Veröffentlicht: (2026)
Resource Allocation in HyperX Networks
von: Cano, Alejandro, et al.
Veröffentlicht: (2026)
von: Cano, Alejandro, et al.
Veröffentlicht: (2026)
DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving
von: Huang, Heyang, et al.
Veröffentlicht: (2025)
von: Huang, Heyang, et al.
Veröffentlicht: (2025)
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
von: Hu, Kan, et al.
Veröffentlicht: (2024)
von: Hu, Kan, et al.
Veröffentlicht: (2024)
Adaptive Resource Allocation for Workflow Containerization on Kubernetes
von: Shan, Chenggang, et al.
Veröffentlicht: (2023)
von: Shan, Chenggang, et al.
Veröffentlicht: (2023)
VDCores: Resource Decoupled Programming and Execution for Asynchronous GPU
von: He, Zijian, et al.
Veröffentlicht: (2026)
von: He, Zijian, et al.
Veröffentlicht: (2026)
Dependency-aware Resource Allocation for Serverless Functions at the Edge
von: Baresi, Luciano, et al.
Veröffentlicht: (2023)
von: Baresi, Luciano, et al.
Veröffentlicht: (2023)
Decentralized Proactive Model Offloading and Resource Allocation for Split and Federated Learning
von: Huang, Binbin, et al.
Veröffentlicht: (2024)
von: Huang, Binbin, et al.
Veröffentlicht: (2024)
Optimizing Resource Allocation and Energy Efficiency in Federated Fog Computing for IoT
von: Shah, Syed Sarmad, et al.
Veröffentlicht: (2025)
von: Shah, Syed Sarmad, et al.
Veröffentlicht: (2025)
Improved Methods of Task Assignment and Resource Allocation with Preemption in Edge Computing Systems
von: Rublein, Caroline, et al.
Veröffentlicht: (2024)
von: Rublein, Caroline, et al.
Veröffentlicht: (2024)
Cloud Resource Allocation with Convex Optimization
von: Boghani, Shayan, et al.
Veröffentlicht: (2025)
von: Boghani, Shayan, et al.
Veröffentlicht: (2025)
Optimizing Hardware Resource Partitioning and Job Allocations on Modern GPUs under Power Caps
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
FedSem: A Resource Allocation Scheme for Federated Learning Assisted Semantic Communication
von: Zhou, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2025)
Resource Allocation of Industry 4.0 Micro-Service Applications across Serverless Fog Federation
von: Hussain, Razin Farhan, et al.
Veröffentlicht: (2024)
von: Hussain, Razin Farhan, et al.
Veröffentlicht: (2024)
Distributed Hierarchical Machine Learning for Joint Resource Allocation and Slice Selection in In-Network Edge Systems
von: Rashid, Sulaiman Muhammad, et al.
Veröffentlicht: (2025)
von: Rashid, Sulaiman Muhammad, et al.
Veröffentlicht: (2025)
AIGC-assisted Federated Learning for Vehicular Edge Intelligence: Vehicle Selection, Resource Allocation and Model Augmentation
von: Qiang, Xianke, et al.
Veröffentlicht: (2025)
von: Qiang, Xianke, et al.
Veröffentlicht: (2025)
Energy-Efficient Joint Offloading and Resource Allocation for Deadline-Constrained Tasks in Multi-Access Edge Computing
von: Gao, Chuanchao, et al.
Veröffentlicht: (2025)
von: Gao, Chuanchao, et al.
Veröffentlicht: (2025)
Workload composition smooths aggregate power demand while sustaining short-horizon ramps in AI data centers
von: Majumder, Subir, et al.
Veröffentlicht: (2026)
von: Majumder, Subir, et al.
Veröffentlicht: (2026)
Cicada: A Pipeline-Efficient Approach to Serverless Inference with Decoupled Management
von: Wu, Z., et al.
Veröffentlicht: (2025)
von: Wu, Z., et al.
Veröffentlicht: (2025)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
DeLIAP e DeLIAJ: Interfaces de biblioteca de Dependabilidade para Python e Julia
von: Irigoyen, Marcos, et al.
Veröffentlicht: (2025)
von: Irigoyen, Marcos, et al.
Veröffentlicht: (2025)
NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding
von: Wang, Jiamin, et al.
Veröffentlicht: (2026)
von: Wang, Jiamin, et al.
Veröffentlicht: (2026)
Resource Allocation and Workload Scheduling for Large-Scale Distributed Deep Learning: A Survey
von: Liang, Feng, et al.
Veröffentlicht: (2024)
von: Liang, Feng, et al.
Veröffentlicht: (2024)
Adaptive, Efficient and Fair Resource Allocation in Cloud Datacenters leveraging Weighted A3C Deep Reinforcement Learning
von: Kumari, Suchi, et al.
Veröffentlicht: (2025)
von: Kumari, Suchi, et al.
Veröffentlicht: (2025)
A Transverse-Read-assisted Fast Valid-Bits Collection in Stochastic Computing MACs for Energy-Efficient in-RTM DNNs
von: Wang, Jihe, et al.
Veröffentlicht: (2024)
von: Wang, Jihe, et al.
Veröffentlicht: (2024)
MultiPaxos Made Complete
von: Liang, Zhiying, et al.
Veröffentlicht: (2024)
von: Liang, Zhiying, et al.
Veröffentlicht: (2024)
The Cost of Garbage Collection for State Machine Replication
von: Liang, Zhiying, et al.
Veröffentlicht: (2024)
von: Liang, Zhiying, et al.
Veröffentlicht: (2024)
CkIO: Parallel File Input for Over-Decomposed Task-Based Systems
von: Jacob, Mathew, et al.
Veröffentlicht: (2024)
von: Jacob, Mathew, et al.
Veröffentlicht: (2024)
Scaling LLM Inference Beyond Amdahl`s Limits via Eliminating Non-Scalable Overheads
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
Resilient Auto-Scaling of Microservice Architectures with Efficient Resource Management
von: Ahmad, Hussain, et al.
Veröffentlicht: (2025)
von: Ahmad, Hussain, et al.
Veröffentlicht: (2025)
ScalePool: Hybrid XLink-CXL Fabric for Composable Resource Disaggregation in Unified Scale-up Domains
von: Woo, Hyein, et al.
Veröffentlicht: (2025)
von: Woo, Hyein, et al.
Veröffentlicht: (2025)
Recursive QAOA for Interference-Aware Resource Allocation in Wireless Networks
von: Chen, Kuan-Cheng, et al.
Veröffentlicht: (2026)
von: Chen, Kuan-Cheng, et al.
Veröffentlicht: (2026)
Graph for Science: From API based Programming to Graph Engine based Programming for HPC
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
Model Partition and Resource Allocation for Split Learning in Vehicular Edge Networks
von: Yu, Lu, et al.
Veröffentlicht: (2024)
von: Yu, Lu, et al.
Veröffentlicht: (2024)
Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation
von: Feng, Weiqi, et al.
Veröffentlicht: (2024)
von: Feng, Weiqi, et al.
Veröffentlicht: (2024)
Collaborative Multi-Agent Reinforcement Learning Approach for Elastic Cloud Resource Scaling
von: Fang, Bruce, et al.
Veröffentlicht: (2025)
von: Fang, Bruce, et al.
Veröffentlicht: (2025)
Autothrottle: A Practical Bi-Level Approach to Resource Management for SLO-Targeted Microservices
von: Wang, Zibo, et al.
Veröffentlicht: (2022)
von: Wang, Zibo, et al.
Veröffentlicht: (2022)
H-EYE: Holistic Resource Modeling and Management for Diversely Scaled Edge-Cloud Systems
von: Dagli, Ismet, et al.
Veröffentlicht: (2024)
von: Dagli, Ismet, et al.
Veröffentlicht: (2024)
Diagonal Scaling: A Multi-Dimensional Resource Model and Optimization Framework for Distributed Databases
von: Abdullah, Shahir, et al.
Veröffentlicht: (2025)
von: Abdullah, Shahir, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
von: Hu, Kan, et al.
Veröffentlicht: (2024) -
DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models
von: Huang, You-Liang, et al.
Veröffentlicht: (2026) -
Resource Allocation in HyperX Networks
von: Cano, Alejandro, et al.
Veröffentlicht: (2026) -
DDiT: Dynamic Resource Allocation for Diffusion Transformer Model Serving
von: Huang, Heyang, et al.
Veröffentlicht: (2025) -
LSRAM: A Lightweight Autoscaling and SLO Resource Allocation Framework for Microservices Based on Gradient Descent
von: Hu, Kan, et al.
Veröffentlicht: (2024)