TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Stojkovic, Jovan, Zhang, Chaojie, Goiri, Íñigo, Choukse, Esha, Qiu, Haoran, Fonseca, Rodrigo, Torrellas, Josep, Bianchini, Ricardo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
von: Qiu, Haoran, et al.
Veröffentlicht: (2026)
von: Qiu, Haoran, et al.
Veröffentlicht: (2026)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Compound AI Systems
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
von: Saurez, Enrique, et al.
Veröffentlicht: (2024)
von: Saurez, Enrique, et al.
Veröffentlicht: (2024)
EcoServe: Designing Carbon-Aware AI Inference Systems
von: Li, Yueying, et al.
Veröffentlicht: (2025)
von: Li, Yueying, et al.
Veröffentlicht: (2025)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
von: Huang, Lexiang, et al.
Veröffentlicht: (2024)
von: Huang, Lexiang, et al.
Veröffentlicht: (2024)
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
von: Iliakopoulou, Nikoleta, et al.
Veröffentlicht: (2024)
von: Iliakopoulou, Nikoleta, et al.
Veröffentlicht: (2024)
Cloud abstractions for AI workloads
von: Canini, Marco, et al.
Veröffentlicht: (2025)
von: Canini, Marco, et al.
Veröffentlicht: (2025)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
von: Jain, Kunal, et al.
Veröffentlicht: (2024)
von: Jain, Kunal, et al.
Veröffentlicht: (2024)
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
von: Oviedo, Felipe, et al.
Veröffentlicht: (2025)
von: Oviedo, Felipe, et al.
Veröffentlicht: (2025)
Towards Cloud Efficiency with Large-scale Workload Characterization
von: Parayil, Anjaly, et al.
Veröffentlicht: (2024)
von: Parayil, Anjaly, et al.
Veröffentlicht: (2024)
Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge Platforms
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2024)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2024)
HEAL: Online Incremental Recovery for Leaderless Distributed Systems Across Persistency Models
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
Towards CXL Resilience to CPU Failures
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
von: Psistakis, Antonis, et al.
Veröffentlicht: (2026)
Designing Datacenter Power Delivery Hierarchies for the AI Era
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
SLO-Aware Scheduling for Large Language Model Inferences
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
von: Raj, Suman, et al.
Veröffentlicht: (2024)
von: Raj, Suman, et al.
Veröffentlicht: (2024)
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
von: Yang, Dingyu, et al.
Veröffentlicht: (2026)
von: Yang, Dingyu, et al.
Veröffentlicht: (2026)
Hotspot-Aware Scheduling of Virtual Machines with Overcommitment for Ultimate Utilization in Cloud Datacenters
von: Wu, Jiaxi, et al.
Veröffentlicht: (2026)
von: Wu, Jiaxi, et al.
Veröffentlicht: (2026)
Power Aware Dynamic Reallocation For Inference
von: Jiang, Yiwei, et al.
Veröffentlicht: (2026)
von: Jiang, Yiwei, et al.
Veröffentlicht: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
von: Wang, Qipeng
Veröffentlicht: (2026)
von: Wang, Qipeng
Veröffentlicht: (2026)
Power-Aware Scheduling for Multi-Center HPC Electricity Cost Optimization
von: Hossain, Abrar, et al.
Veröffentlicht: (2025)
von: Hossain, Abrar, et al.
Veröffentlicht: (2025)
Power Aware Container Placement in Cloud Computing with Affinity and Cubic Power Model
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2024)
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2024)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
von: Wolfrath, Joel, et al.
Veröffentlicht: (2025)
von: Wolfrath, Joel, et al.
Veröffentlicht: (2025)
The AI_INFN Platform: Artificial Intelligence Development in the Cloud
von: Anderlini, Lucio, et al.
Veröffentlicht: (2025)
von: Anderlini, Lucio, et al.
Veröffentlicht: (2025)
Nanvix: A Multikernel OS Design for High-Density Serverless Deployments
von: Segarra, Carlos, et al.
Veröffentlicht: (2026)
von: Segarra, Carlos, et al.
Veröffentlicht: (2026)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
ACE-GNN: Adaptive GNN Co-Inference with System-Aware Scheduling in Dynamic Edge Environments
von: Zhou, Ao, et al.
Veröffentlicht: (2025)
von: Zhou, Ao, et al.
Veröffentlicht: (2025)
An Elastic Job Scheduler for HPC Applications on the Cloud
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
Efficient Probabilistic Workflow Scheduling for IaaS Clouds
von: Russo, Gabriele Russo, et al.
Veröffentlicht: (2024)
von: Russo, Gabriele Russo, et al.
Veröffentlicht: (2024)
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
von: Chang, Tzu-Tao, et al.
Veröffentlicht: (2025)
von: Chang, Tzu-Tao, et al.
Veröffentlicht: (2025)
e112: A Context-Aware Mobile Emergency Communication Platform Leveraging Smartphone Sensing and Cloud Services
von: Ioannidou, Katerina, et al.
Veröffentlicht: (2026)
von: Ioannidou, Katerina, et al.
Veröffentlicht: (2026)
GeoFaaS: An Edge-to-Cloud FaaS Platform
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Malekabbasi, Mohammadreza, et al.
Veröffentlicht: (2024)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
von: Qi, S., et al.
Veröffentlicht: (2024)
von: Qi, S., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024) -
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024) -
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
von: Qiu, Haoran, et al.
Veröffentlicht: (2026) -
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025) -
Towards Resource-Efficient Compound AI Systems
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)