Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wilkins, Grant, Keshav, Srinivasan, Mortier, Richard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
Towards Carbon-Aware Container Orchestration: Predicting Workload Energy Consumption with Federated Learning
von: Saad, Zainab, et al.
Veröffentlicht: (2025)
von: Saad, Zainab, et al.
Veröffentlicht: (2025)
Robust Synchronisation for Federated Learning in The Face of Correlated Device Failure
von: Behfar, Stefan, et al.
Veröffentlicht: (2026)
von: Behfar, Stefan, et al.
Veröffentlicht: (2026)
Hybrid Learning and Optimization-Based Dynamic Scheduling for DL Workloads on Heterogeneous GPU Clusters
von: Dongare, Shruti, et al.
Veröffentlicht: (2025)
von: Dongare, Shruti, et al.
Veröffentlicht: (2025)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
Quantifying Energy and Cost Benefits of Hybrid Edge Cloud: Analysis of Traditional and Agentic Workloads
von: Alamouti, Siavash
Veröffentlicht: (2025)
von: Alamouti, Siavash
Veröffentlicht: (2025)
High-Throughput LLM inference on Heterogeneous Clusters
von: Xiong, Yi, et al.
Veröffentlicht: (2025)
von: Xiong, Yi, et al.
Veröffentlicht: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
von: Zhang, Ping, et al.
Veröffentlicht: (2024)
von: Zhang, Ping, et al.
Veröffentlicht: (2024)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
von: Luo, Xinhao, et al.
Veröffentlicht: (2025)
von: Luo, Xinhao, et al.
Veröffentlicht: (2025)
HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads
von: Agyemang, Justice Owusu, et al.
Veröffentlicht: (2026)
von: Agyemang, Justice Owusu, et al.
Veröffentlicht: (2026)
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
Designing Datacenter Power Delivery Hierarchies for the AI Era
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
Duration-Informed Workload Scheduler
von: Loreti, Daniela, et al.
Veröffentlicht: (2026)
von: Loreti, Daniela, et al.
Veröffentlicht: (2026)
Workload Schedulers -- Genesis, Algorithms and Differences
von: Sliwko, Leszek, et al.
Veröffentlicht: (2025)
von: Sliwko, Leszek, et al.
Veröffentlicht: (2025)
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
von: Liu, Ziyue, et al.
Veröffentlicht: (2026)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
Tesserae: Scalable Placement Policies for Deep Learning Workloads
von: Bian, Song, et al.
Veröffentlicht: (2025)
von: Bian, Song, et al.
Veröffentlicht: (2025)
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models
von: Xu, Si, et al.
Veröffentlicht: (2024)
von: Xu, Si, et al.
Veröffentlicht: (2024)
Failure-Resilient Distributed Inference with Model Compression over Heterogeneous Edge Devices
von: Wang, Li, et al.
Veröffentlicht: (2024)
von: Wang, Li, et al.
Veröffentlicht: (2024)
Accelerating LLM Inference with Precomputed Query Storage
von: Park, Jay H., et al.
Veröffentlicht: (2025)
von: Park, Jay H., et al.
Veröffentlicht: (2025)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
Decentralized AI: Permissionless LLM Inference on POKT Network
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
von: Chen, Le, et al.
Veröffentlicht: (2025)
von: Chen, Le, et al.
Veröffentlicht: (2025)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
von: Özcan, Miray, et al.
Veröffentlicht: (2025)
von: Özcan, Miray, et al.
Veröffentlicht: (2025)
Distributed Inference on Mobile Edge and Cloud: A Data-Cartography based Clustering Approach
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2024)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
Resource Allocation and Workload Scheduling for Large-Scale Distributed Deep Learning: A Survey
von: Liang, Feng, et al.
Veröffentlicht: (2024)
von: Liang, Feng, et al.
Veröffentlicht: (2024)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
von: Chandrasekar, Ashok, et al.
Veröffentlicht: (2026)
von: Chandrasekar, Ashok, et al.
Veröffentlicht: (2026)
Seesaw: High-throughput LLM Inference via Model Re-sharding
von: Su, Qidong, et al.
Veröffentlicht: (2025)
von: Su, Qidong, et al.
Veröffentlicht: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
von: Wilkins, Grant, et al.
Veröffentlicht: (2024) -
Towards Carbon-Aware Container Orchestration: Predicting Workload Energy Consumption with Federated Learning
von: Saad, Zainab, et al.
Veröffentlicht: (2025) -
Robust Synchronisation for Federated Learning in The Face of Correlated Device Failure
von: Behfar, Stefan, et al.
Veröffentlicht: (2026) -
Hybrid Learning and Optimization-Based Dynamic Scheduling for DL Workloads on Heterogeneous GPU Clusters
von: Dongare, Shruti, et al.
Veröffentlicht: (2025) -
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)