DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
Fuente:
arXiv
Salvato in:
| Autori principali: | Stojkovic, Jovan, Zhang, Chaojie, Goiri, Íñigo, Torrellas, Josep, Choukse, Esha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
di: Iliakopoulou, Nikoleta, et al.
Pubblicazione: (2024)
di: Iliakopoulou, Nikoleta, et al.
Pubblicazione: (2024)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
di: DeBole, Michael V., et al.
Pubblicazione: (2025)
di: DeBole, Michael V., et al.
Pubblicazione: (2025)
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
di: Ortega, Cristobal, et al.
Pubblicazione: (2024)
di: Ortega, Cristobal, et al.
Pubblicazione: (2024)
NPU Design for Diffusion Language Model Inference
di: Lou, Binglei, et al.
Pubblicazione: (2026)
di: Lou, Binglei, et al.
Pubblicazione: (2026)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2025)
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2025)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
di: Kakolyris, Andreas Kosmas, et al.
Pubblicazione: (2024)
di: Kakolyris, Andreas Kosmas, et al.
Pubblicazione: (2024)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
di: Renney, Harri, et al.
Pubblicazione: (2026)
di: Renney, Harri, et al.
Pubblicazione: (2026)
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2025)
di: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Pubblicazione: (2025)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2024)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
di: Liu, Yiqi, et al.
Pubblicazione: (2026)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
di: Kubwimana, Benjamin, et al.
Pubblicazione: (2025)
di: Kubwimana, Benjamin, et al.
Pubblicazione: (2025)
DCO: Dynamic Cache Orchestration for LLM Accelerators through Predictive Management
di: Zhou, Zhongchun, et al.
Pubblicazione: (2025)
di: Zhou, Zhongchun, et al.
Pubblicazione: (2025)
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
di: Yüzügüler, Ahmet Caner, et al.
Pubblicazione: (2025)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
di: Zhu, Wenbin, et al.
Pubblicazione: (2025)
di: Zhu, Wenbin, et al.
Pubblicazione: (2025)
Improving AI Efficiency in Data Centres by Power Dynamic Response
di: Marinoni, Andrea, et al.
Pubblicazione: (2025)
di: Marinoni, Andrea, et al.
Pubblicazione: (2025)
Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference
di: Zhao, Yiren, et al.
Pubblicazione: (2026)
di: Zhao, Yiren, et al.
Pubblicazione: (2026)
Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective
di: Shi, Tianyao, et al.
Pubblicazione: (2025)
di: Shi, Tianyao, et al.
Pubblicazione: (2025)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
di: Chung, Euijun, et al.
Pubblicazione: (2026)
di: Chung, Euijun, et al.
Pubblicazione: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
di: Meng, William, et al.
Pubblicazione: (2025)
di: Meng, William, et al.
Pubblicazione: (2025)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
di: Qiu, Haoran, et al.
Pubblicazione: (2026)
di: Qiu, Haoran, et al.
Pubblicazione: (2026)
Power Stabilization for AI Training Datacenters
di: Choukse, Esha, et al.
Pubblicazione: (2025)
di: Choukse, Esha, et al.
Pubblicazione: (2025)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
di: Li, Jiaxi, et al.
Pubblicazione: (2025)
di: Li, Jiaxi, et al.
Pubblicazione: (2025)
Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference
di: Yu, Zhongkai, et al.
Pubblicazione: (2025)
di: Yu, Zhongkai, et al.
Pubblicazione: (2025)
TriMoE: Augmenting GPU with AMX-Enabled CPU and DIMM-NDP for High-Throughput MoE Inference via Offloading
di: Pan, Yudong, et al.
Pubblicazione: (2026)
di: Pan, Yudong, et al.
Pubblicazione: (2026)
HPU: High-Bandwidth Processing Unit for Scalable, Cost-effective LLM Inference via GPU Co-processing
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
di: Rhee, Myunghyun, et al.
Pubblicazione: (2025)
The DMA Streaming Framework: Kernel-Level Buffer Orchestration for High-Performance AI Data Paths
di: Graziano, Marco
Pubblicazione: (2026)
di: Graziano, Marco
Pubblicazione: (2026)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
di: Lin, Bin, et al.
Pubblicazione: (2024)
di: Lin, Bin, et al.
Pubblicazione: (2024)
Advancing AI-assisted Hardware Design with Hierarchical Decentralized Training and Personalized Inference-Time Optimization
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
di: Chen, Hao Mark, et al.
Pubblicazione: (2025)
Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving
di: Zhao, Juntao, et al.
Pubblicazione: (2025)
di: Zhao, Juntao, et al.
Pubblicazione: (2025)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
di: Muntaka, Siddique Abubakr, et al.
Pubblicazione: (2026)
di: Muntaka, Siddique Abubakr, et al.
Pubblicazione: (2026)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
di: Lei, Jianlong, et al.
Pubblicazione: (2026)
di: Lei, Jianlong, et al.
Pubblicazione: (2026)
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
di: Yeo, Gwangoo, et al.
Pubblicazione: (2024)
di: Yeo, Gwangoo, et al.
Pubblicazione: (2024)
Intent-Driven Storage Systems: From Low-Level Tuning to High-Level Understanding
di: Bergman, Shai, et al.
Pubblicazione: (2025)
di: Bergman, Shai, et al.
Pubblicazione: (2025)
PiKV: KV Cache Management System for Mixture of Experts
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
di: Zhao, Dan, et al.
Pubblicazione: (2024)
di: Zhao, Dan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024) -
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025) -
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025) -
Splitwise: Efficient generative LLM inference using phase splitting
di: Patel, Pratyush, et al.
Pubblicazione: (2023) -
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
di: Iliakopoulou, Nikoleta, et al.
Pubblicazione: (2024)