From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
Fuente:
arXiv
Salvato in:
| Autori principali: | Wilkins, Grant, Kazhamiaka, Fiodar, Rajagopal, Ram |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Designing Datacenter Power Delivery Hierarchies for the AI Era
di: Wilkins, Grant, et al.
Pubblicazione: (2026)
di: Wilkins, Grant, et al.
Pubblicazione: (2026)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
di: Oviedo, Felipe, et al.
Pubblicazione: (2025)
di: Oviedo, Felipe, et al.
Pubblicazione: (2025)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
di: Zhang, Hongbin, et al.
Pubblicazione: (2026)
di: Zhang, Hongbin, et al.
Pubblicazione: (2026)
BSODiag: A Global Diagnosis Framework for Batch Servers Outage in Large-scale Cloud Infrastructure Systems
di: Duan, Tao, et al.
Pubblicazione: (2025)
di: Duan, Tao, et al.
Pubblicazione: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
di: Arya, Mayank, et al.
Pubblicazione: (2025)
di: Arya, Mayank, et al.
Pubblicazione: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025)
di: Karfakis, George, et al.
Pubblicazione: (2025)
To Compress or Not To Compress: Energy Trade-Offs and Benefits of Lossy Compressed I/O
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
di: Zhang, Yuning, et al.
Pubblicazione: (2025)
di: Zhang, Yuning, et al.
Pubblicazione: (2025)
Are Bus-Mounted Edge Servers Feasible?
di: Li, Xuezhi, et al.
Pubblicazione: (2025)
di: Li, Xuezhi, et al.
Pubblicazione: (2025)
Node Compass: Multilevel Tracing and Debugging of Request Executions in JavaScript-Based Web-Servers
di: Kabamba, Herve Mbikayi, et al.
Pubblicazione: (2023)
di: Kabamba, Herve Mbikayi, et al.
Pubblicazione: (2023)
VQ-LLM: High-performance Code Generation for Vector Quantization Augmented LLM Inference
di: Liu, Zihan, et al.
Pubblicazione: (2025)
di: Liu, Zihan, et al.
Pubblicazione: (2025)
Practical Federated Learning without a Server
di: Dhasade, Akash, et al.
Pubblicazione: (2025)
di: Dhasade, Akash, et al.
Pubblicazione: (2025)
Experimental Analysis of Server-Side Caching for Web Performance
di: Umar, Mohammad, et al.
Pubblicazione: (2026)
di: Umar, Mohammad, et al.
Pubblicazione: (2026)
SCARIF: Towards Carbon Modeling of Cloud Servers with Accelerators
di: Ji, Shixin, et al.
Pubblicazione: (2024)
di: Ji, Shixin, et al.
Pubblicazione: (2024)
Revisiting Parameter Server in LLM Post-Training
di: Wan, Xinyi, et al.
Pubblicazione: (2026)
di: Wan, Xinyi, et al.
Pubblicazione: (2026)
Power Aware Dynamic Reallocation For Inference
di: Jiang, Yiwei, et al.
Pubblicazione: (2026)
di: Jiang, Yiwei, et al.
Pubblicazione: (2026)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
di: Chen, Jiu, et al.
Pubblicazione: (2026)
di: Chen, Jiu, et al.
Pubblicazione: (2026)
Analysis of Server Throughput For Managed Big Data Analytics Frameworks
di: Anagnostakis, Emmanouil, et al.
Pubblicazione: (2025)
di: Anagnostakis, Emmanouil, et al.
Pubblicazione: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
Experiences with Model Context Protocol Servers for Science and High Performance Computing
di: Pan, Haochen, et al.
Pubblicazione: (2025)
di: Pan, Haochen, et al.
Pubblicazione: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
di: Li, Suyi, et al.
Pubblicazione: (2024)
di: Li, Suyi, et al.
Pubblicazione: (2024)
SECO: Secure Inference With Model Splitting Across Multi-Server Hierarchy
di: Chen, Shuangyi, et al.
Pubblicazione: (2024)
di: Chen, Shuangyi, et al.
Pubblicazione: (2024)
Federated Inference for Heterogeneous LLM Communication and Collaboration
di: Chen, Zihan, et al.
Pubblicazione: (2026)
di: Chen, Zihan, et al.
Pubblicazione: (2026)
Cloud Native System for LLM Inference Serving
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
Enabling Dynamic Sparsity in Quantized LLM Inference
di: Wang, Rongxiang, et al.
Pubblicazione: (2025)
di: Wang, Rongxiang, et al.
Pubblicazione: (2025)
SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
di: Xiong, Yifan, et al.
Pubblicazione: (2024)
di: Xiong, Yifan, et al.
Pubblicazione: (2024)
FedSZ: Leveraging Error-Bounded Lossy Compression for Federated Learning Communications
di: Wilkins, Grant, et al.
Pubblicazione: (2023)
di: Wilkins, Grant, et al.
Pubblicazione: (2023)
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
di: Ding, Eric, et al.
Pubblicazione: (2026)
di: Ding, Eric, et al.
Pubblicazione: (2026)
WANSpec: Leveraging Global Compute Capacity for LLM Inference
di: Martin, Noah, et al.
Pubblicazione: (2026)
di: Martin, Noah, et al.
Pubblicazione: (2026)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
di: Xu, Chuhao, et al.
Pubblicazione: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
di: Wu, Panlong, et al.
Pubblicazione: (2025)
di: Wu, Panlong, et al.
Pubblicazione: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
di: Rajashekar, Kolichala, et al.
Pubblicazione: (2025)
di: Rajashekar, Kolichala, et al.
Pubblicazione: (2025)
Distributed On-Device LLM Inference With Over-the-Air Computation
di: Zhang, Kai, et al.
Pubblicazione: (2025)
di: Zhang, Kai, et al.
Pubblicazione: (2025)
Efficient LLM Inference with Activation Checkpointing and Hybrid Caching
di: Lee, Sanghyeon, et al.
Pubblicazione: (2025)
di: Lee, Sanghyeon, et al.
Pubblicazione: (2025)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
TailBench++: Flexible Multi-Client, Multi-Server Benchmarking for Latency-Critical Workloads
di: Li, Zhilin, et al.
Pubblicazione: (2025)
di: Li, Zhilin, et al.
Pubblicazione: (2025)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
di: Da, Wei, et al.
Pubblicazione: (2026)
di: Da, Wei, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Designing Datacenter Power Delivery Hierarchies for the AI Era
di: Wilkins, Grant, et al.
Pubblicazione: (2026) -
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
di: Wilkins, Grant, et al.
Pubblicazione: (2024) -
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
di: Oviedo, Felipe, et al.
Pubblicazione: (2025) -
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025) -
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
di: Wilkins, Grant, et al.
Pubblicazione: (2024)