Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
Fuente:
arXiv
Salvato in:
| Autori principali: | Stojkovic, Jovan, Zhang, Chaojie, Goiri, Íñigo, Bianchini, Ricardo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
Power Stabilization for AI Training Datacenters
di: Choukse, Esha, et al.
Pubblicazione: (2025)
di: Choukse, Esha, et al.
Pubblicazione: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
The DMA Streaming Framework: Kernel-Level Buffer Orchestration for High-Performance AI Data Paths
di: Graziano, Marco
Pubblicazione: (2026)
di: Graziano, Marco
Pubblicazione: (2026)
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
di: Wang, Jing, et al.
Pubblicazione: (2025)
di: Wang, Jing, et al.
Pubblicazione: (2025)
Intent-Driven Storage Systems: From Low-Level Tuning to High-Level Understanding
di: Bergman, Shai, et al.
Pubblicazione: (2025)
di: Bergman, Shai, et al.
Pubblicazione: (2025)
HyperOffload: Graph-Driven Hierarchical Memory Management for Large Language Models on SuperNode Architectures
di: Liu, Fangxin, et al.
Pubblicazione: (2026)
di: Liu, Fangxin, et al.
Pubblicazione: (2026)
ODIN-Based CPU-GPU Architecture with Replay-Driven Simulation and Emulation
di: Dorairaj, Nij, et al.
Pubblicazione: (2026)
di: Dorairaj, Nij, et al.
Pubblicazione: (2026)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
di: Chu, Xiaoyu, et al.
Pubblicazione: (2024)
di: Chu, Xiaoyu, et al.
Pubblicazione: (2024)
Improving AI Efficiency in Data Centres by Power Dynamic Response
di: Marinoni, Andrea, et al.
Pubblicazione: (2025)
di: Marinoni, Andrea, et al.
Pubblicazione: (2025)
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
di: Zhao, Yiren, et al.
Pubblicazione: (2026)
di: Zhao, Yiren, et al.
Pubblicazione: (2026)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
di: Zhao, Dan, et al.
Pubblicazione: (2024)
di: Zhao, Dan, et al.
Pubblicazione: (2024)
Debunking the CUDA Myth Towards GPU-based AI Systems
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
di: Lee, Yunjae, et al.
Pubblicazione: (2024)
Modernizing Amdahl's Law: How AI Scaling Laws Shape Computer Architecture
di: Lu, Chien-Ping
Pubblicazione: (2026)
di: Lu, Chien-Ping
Pubblicazione: (2026)
Efficient Edge AI: Deploying Convolutional Neural Networks on FPGA with the Gemmini Accelerator
di: Peccia, Federico Nicolas, et al.
Pubblicazione: (2024)
di: Peccia, Federico Nicolas, et al.
Pubblicazione: (2024)
Exploring energy consumption of AI frameworks on a 64-core RV64 Server CPU
di: Malenza, Giulio, et al.
Pubblicazione: (2025)
di: Malenza, Giulio, et al.
Pubblicazione: (2025)
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
di: Canakci, Burcu, et al.
Pubblicazione: (2025)
di: Canakci, Burcu, et al.
Pubblicazione: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
RedFuser: An Automatic Operator Fusion Framework for Cascaded Reductions on AI Accelerators
di: Tang, Xinsheng, et al.
Pubblicazione: (2026)
di: Tang, Xinsheng, et al.
Pubblicazione: (2026)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
di: Zhou, Zhongchun, et al.
Pubblicazione: (2025)
di: Zhou, Zhongchun, et al.
Pubblicazione: (2025)
Investigating Memory Failure Prediction Across CPU Architectures
di: Yu, Qiao, et al.
Pubblicazione: (2024)
di: Yu, Qiao, et al.
Pubblicazione: (2024)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
di: DeBole, Michael V., et al.
Pubblicazione: (2025)
di: DeBole, Michael V., et al.
Pubblicazione: (2025)
Designing Datacenter Power Delivery Hierarchies for the AI Era
di: Wilkins, Grant, et al.
Pubblicazione: (2026)
di: Wilkins, Grant, et al.
Pubblicazione: (2026)
PiKV: KV Cache Management System for Mixture of Experts
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
di: Kubwimana, Benjamin, et al.
Pubblicazione: (2025)
di: Kubwimana, Benjamin, et al.
Pubblicazione: (2025)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
di: Zhu, Wenbin, et al.
Pubblicazione: (2025)
di: Zhu, Wenbin, et al.
Pubblicazione: (2025)
SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
di: Li, Jonathan, et al.
Pubblicazione: (2025)
di: Li, Jonathan, et al.
Pubblicazione: (2025)
Taming Asynchronous CPU-GPU Coupling for Frequency-aware Latency Estimation on Mobile Edge
di: Chen, Jiesong, et al.
Pubblicazione: (2026)
di: Chen, Jiesong, et al.
Pubblicazione: (2026)
Co-design of a novel CMOS highly parallel, low-power, multi-chip neural network accelerator
di: Hokenmaier, W, et al.
Pubblicazione: (2024)
di: Hokenmaier, W, et al.
Pubblicazione: (2024)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
di: Pan, Yudong, et al.
Pubblicazione: (2026)
di: Pan, Yudong, et al.
Pubblicazione: (2026)
Strict Partitioning for Sporadic Rigid Gang Tasks
di: Sun, Binqi, et al.
Pubblicazione: (2024)
di: Sun, Binqi, et al.
Pubblicazione: (2024)
NPU Design for Diffusion Language Model Inference
di: Lou, Binglei, et al.
Pubblicazione: (2026)
di: Lou, Binglei, et al.
Pubblicazione: (2026)
Forge-UGC: FX optimization and register-graph engine for universal graph compiler
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
PhD Thesis Summary: Methods for Reliability Assessment and Enhancement of Deep Neural Network Hardware Accelerators
di: Taheri, Mahdi
Pubblicazione: (2026)
di: Taheri, Mahdi
Pubblicazione: (2026)
COMET: Neural Cost Model Explanation Framework
di: Chaudhary, Isha, et al.
Pubblicazione: (2023)
di: Chaudhary, Isha, et al.
Pubblicazione: (2023)
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
di: Kumar, Deepak, et al.
Pubblicazione: (2025)
di: Kumar, Deepak, et al.
Pubblicazione: (2025)
HLS4PC: A Parametrizable Framework For Accelerating Point-Based 3D Point Cloud Models on FPGA
di: Pal, Amur Saqib, et al.
Pubblicazione: (2025)
di: Pal, Amur Saqib, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024) -
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024) -
Power Stabilization for AI Training Datacenters
di: Choukse, Esha, et al.
Pubblicazione: (2025) -
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025) -
The DMA Streaming Framework: Kernel-Level Buffer Orchestration for High-Performance AI Data Paths
di: Graziano, Marco
Pubblicazione: (2026)