GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Qunyou, Huang, Darong, Zapater, Marina, Atienza, David |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
di: Liu, Qunyou, et al.
Pubblicazione: (2025)
di: Liu, Qunyou, et al.
Pubblicazione: (2025)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
di: Lysenstøen, Christian
Pubblicazione: (2026)
di: Lysenstøen, Christian
Pubblicazione: (2026)
MatrixFlow: System-Accelerator co-design for high-performance transformer applications
di: Liu, Qunyou, et al.
Pubblicazione: (2025)
di: Liu, Qunyou, et al.
Pubblicazione: (2025)
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design
di: Liu, Qunyou, et al.
Pubblicazione: (2026)
di: Liu, Qunyou, et al.
Pubblicazione: (2026)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
di: Liu, Qunyou, et al.
Pubblicazione: (2026)
di: Liu, Qunyou, et al.
Pubblicazione: (2026)
CloudFormer: An Attention-based Performance Prediction for Public Clouds with Unknown Workload
di: Shahbazinia, Amirhossein, et al.
Pubblicazione: (2025)
di: Shahbazinia, Amirhossein, et al.
Pubblicazione: (2025)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
di: Ma, Xinyue, et al.
Pubblicazione: (2026)
di: Ma, Xinyue, et al.
Pubblicazione: (2026)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
di: Zhang, Quqing, et al.
Pubblicazione: (2026)
di: Zhang, Quqing, et al.
Pubblicazione: (2026)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
di: Kakolyris, Andreas Kosmas, et al.
Pubblicazione: (2024)
di: Kakolyris, Andreas Kosmas, et al.
Pubblicazione: (2024)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
di: Kong, Linghao, et al.
Pubblicazione: (2026)
di: Kong, Linghao, et al.
Pubblicazione: (2026)
Prompt-Aware Scheduling for Low-Latency LLM Serving
di: Tao, Yiheng, et al.
Pubblicazione: (2025)
di: Tao, Yiheng, et al.
Pubblicazione: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
di: Fan, Ruibo, et al.
Pubblicazione: (2026)
di: Fan, Ruibo, et al.
Pubblicazione: (2026)
Reconsidering the performance of DEVS modeling and simulation environments using the DEVStone benchmark
di: Risco-Martín, José L., et al.
Pubblicazione: (2024)
di: Risco-Martín, José L., et al.
Pubblicazione: (2024)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
di: Liu, Xiaoxuan, et al.
Pubblicazione: (2024)
di: Liu, Xiaoxuan, et al.
Pubblicazione: (2024)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
di: Qi, S., et al.
Pubblicazione: (2024)
di: Qi, S., et al.
Pubblicazione: (2024)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
di: Shi, Tianyao, et al.
Pubblicazione: (2024)
di: Shi, Tianyao, et al.
Pubblicazione: (2024)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
di: Yu, Shan, et al.
Pubblicazione: (2025)
di: Yu, Shan, et al.
Pubblicazione: (2025)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
di: He, Jiaao, et al.
Pubblicazione: (2024)
di: He, Jiaao, et al.
Pubblicazione: (2024)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
di: Jayakody, Shakya, et al.
Pubblicazione: (2026)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
di: Yang, Shang, et al.
Pubblicazione: (2025)
di: Yang, Shang, et al.
Pubblicazione: (2025)
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
di: Agullo, Ferran, et al.
Pubblicazione: (2025)
di: Agullo, Ferran, et al.
Pubblicazione: (2025)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
di: Liu, Hongyao, et al.
Pubblicazione: (2026)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
di: Wang, Yuxin, et al.
Pubblicazione: (2024)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
di: Lin, Yujun, et al.
Pubblicazione: (2024)
di: Lin, Yujun, et al.
Pubblicazione: (2024)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
di: Yi, Qingao, et al.
Pubblicazione: (2025)
di: Yi, Qingao, et al.
Pubblicazione: (2025)
Quantifying the Generalization Gap in Seizure Detection: A Large-Scale Empirical Benchmark via the SzCORE Challenge
di: Dan, Jonathan, et al.
Pubblicazione: (2025)
di: Dan, Jonathan, et al.
Pubblicazione: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
di: Karfakis, George, et al.
Pubblicazione: (2025)
di: Karfakis, George, et al.
Pubblicazione: (2025)
SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference
di: Cavagna, Hiari Pizzini, et al.
Pubblicazione: (2026)
di: Cavagna, Hiari Pizzini, et al.
Pubblicazione: (2026)
Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey
di: Moslem, Yasmin, et al.
Pubblicazione: (2026)
di: Moslem, Yasmin, et al.
Pubblicazione: (2026)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
di: Yang, Hanmei, et al.
Pubblicazione: (2024)
di: Yang, Hanmei, et al.
Pubblicazione: (2024)
Benchmarking Dynamic SLO Compliance in Distributed Computing Continuum Systems
di: Lapkovskis, Alfreds, et al.
Pubblicazione: (2025)
di: Lapkovskis, Alfreds, et al.
Pubblicazione: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
di: Fu, Zizhuo, et al.
Pubblicazione: (2025)
di: Fu, Zizhuo, et al.
Pubblicazione: (2025)
LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
di: Liu, Jiashuo, et al.
Pubblicazione: (2025)
di: Liu, Jiashuo, et al.
Pubblicazione: (2025)
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
di: Yuan, Ziqi, et al.
Pubblicazione: (2025)
di: Yuan, Ziqi, et al.
Pubblicazione: (2025)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
di: Jiang, Chaoyi, et al.
Pubblicazione: (2024)
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
di: Ray, Kaustabha, et al.
Pubblicazione: (2025)
di: Ray, Kaustabha, et al.
Pubblicazione: (2025)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
di: Sunesh, Aman, et al.
Pubblicazione: (2026)
di: Sunesh, Aman, et al.
Pubblicazione: (2026)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
di: Xue, Leyang, et al.
Pubblicazione: (2024)
di: Xue, Leyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
di: Liu, Qunyou, et al.
Pubblicazione: (2025) -
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
di: Ziller, Thomas, et al.
Pubblicazione: (2026) -
SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving
di: Lysenstøen, Christian
Pubblicazione: (2026) -
MatrixFlow: System-Accelerator co-design for high-performance transformer applications
di: Liu, Qunyou, et al.
Pubblicazione: (2025) -
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design
di: Liu, Qunyou, et al.
Pubblicazione: (2026)