GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ziller, Thomas, Ilager, Shashikant, Tundo, Alessandro, Bartocci, Ezio, Mariani, Leonardo, Brandic, Ivona |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
DynaSplit: A Hardware-Software Co-Design Framework for Energy-Aware Inference on Edge
par: May, Daniel, et autres
Publié: (2024)
par: May, Daniel, et autres
Publié: (2024)
GREEN-CODE: Learning to Optimize Energy Efficiency in LLM-based Code Generation
par: Ilager, Shashikant, et autres
Publié: (2025)
par: Ilager, Shashikant, et autres
Publié: (2025)
A Decentralized and Self-Adaptive Approach for Monitoring Volatile Edge Environments
par: Ilager, Shashikant, et autres
Publié: (2024)
par: Ilager, Shashikant, et autres
Publié: (2024)
Characterizing LLM Inference Energy-Performance Tradeoffs across Workloads and GPU Scaling
par: Maliakel, Paul Joe, et autres
Publié: (2025)
par: Maliakel, Paul Joe, et autres
Publié: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
par: Lei, Jianlong, et autres
Publié: (2026)
par: Lei, Jianlong, et autres
Publié: (2026)
ABBA-VSM: Time Series Classification using Symbolic Representation on the Edge
par: Kanatbekova, Meerzhan, et autres
Publié: (2024)
par: Kanatbekova, Meerzhan, et autres
Publié: (2024)
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
par: Liu, Qunyou, et autres
Publié: (2025)
par: Liu, Qunyou, et autres
Publié: (2025)
FLIGAN: Enhancing Federated Learning with Incomplete Data using GAN
par: Maliakel, Paul Joe, et autres
Publié: (2024)
par: Maliakel, Paul Joe, et autres
Publié: (2024)
Leveraging LLMs for Structured Information Extraction and Analysis from Cloud Incident Reports (Work In Progress Paper)
par: Chu, Xiaoyu, et autres
Publié: (2026)
par: Chu, Xiaoyu, et autres
Publié: (2026)
FRESCO: Fast and Reliable Edge Offloading with Reputation-based Hybrid Smart Contracts
par: Zilic, Josip, et autres
Publié: (2024)
par: Zilic, Josip, et autres
Publié: (2024)
Breaking Down Quantum Compilation: Profiling and Identifying Costly Passes
par: Zilk, Felix, et autres
Publié: (2025)
par: Zilk, Felix, et autres
Publié: (2025)
Dynamic Model Routing and Cascading for Efficient LLM Inference: A Survey
par: Moslem, Yasmin, et autres
Publié: (2026)
par: Moslem, Yasmin, et autres
Publié: (2026)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
par: Fan, Ruibo, et autres
Publié: (2026)
par: Fan, Ruibo, et autres
Publié: (2026)
INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
par: Šabanović, Ahmed, et autres
Publié: (2026)
par: Šabanović, Ahmed, et autres
Publié: (2026)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
par: Zhang, Quqing, et autres
Publié: (2026)
par: Zhang, Quqing, et autres
Publié: (2026)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
par: Fu, Zizhuo, et autres
Publié: (2025)
par: Fu, Zizhuo, et autres
Publié: (2025)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
par: Ma, Xinyue, et autres
Publié: (2026)
par: Ma, Xinyue, et autres
Publié: (2026)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
par: Yu, Shan, et autres
Publié: (2025)
par: Yu, Shan, et autres
Publié: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
par: Kong, Linghao, et autres
Publié: (2026)
par: Kong, Linghao, et autres
Publié: (2026)
Prompt-Aware Scheduling for Low-Latency LLM Serving
par: Tao, Yiheng, et autres
Publié: (2025)
par: Tao, Yiheng, et autres
Publié: (2025)
SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference
par: Cavagna, Hiari Pizzini, et autres
Publié: (2026)
par: Cavagna, Hiari Pizzini, et autres
Publié: (2026)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
par: Chu, Xiaoyu, et autres
Publié: (2024)
par: Chu, Xiaoyu, et autres
Publié: (2024)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
par: Liu, Hongyao, et autres
Publié: (2026)
par: Liu, Hongyao, et autres
Publié: (2026)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
par: Suo, Jiashun, et autres
Publié: (2025)
par: Suo, Jiashun, et autres
Publié: (2025)
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
par: Ray, Kaustabha, et autres
Publié: (2025)
par: Ray, Kaustabha, et autres
Publié: (2025)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
par: Jiang, Chaoyi, et autres
Publié: (2024)
par: Jiang, Chaoyi, et autres
Publié: (2024)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
par: Karfakis, George, et autres
Publié: (2025)
par: Karfakis, George, et autres
Publié: (2025)
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
par: Dong, Ximing, et autres
Publié: (2026)
par: Dong, Ximing, et autres
Publié: (2026)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
par: Liu, Xiaoxuan, et autres
Publié: (2024)
par: Liu, Xiaoxuan, et autres
Publié: (2024)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
par: Ding, Jiabiao, et autres
Publié: (2026)
par: Ding, Jiabiao, et autres
Publié: (2026)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
par: Dutt, Anurag, et autres
Publié: (2025)
par: Dutt, Anurag, et autres
Publié: (2025)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
par: He, Jiaao, et autres
Publié: (2024)
par: He, Jiaao, et autres
Publié: (2024)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
par: Yao, Feiyu, et autres
Publié: (2026)
par: Yao, Feiyu, et autres
Publié: (2026)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
par: Xue, Leyang, et autres
Publié: (2024)
par: Xue, Leyang, et autres
Publié: (2024)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
par: Shin, Jiho, et autres
Publié: (2024)
par: Shin, Jiho, et autres
Publié: (2024)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
par: Jayakody, Shakya, et autres
Publié: (2026)
par: Jayakody, Shakya, et autres
Publié: (2026)
It's Not Easy Being Green: On the Energy Efficiency of Programming Languages
par: van Kempen, Nicolas, et autres
Publié: (2024)
par: van Kempen, Nicolas, et autres
Publié: (2024)
Energy-Efficient Transformer Inference: Optimization Strategies for Time Series Classification
par: Kermani, Arshia, et autres
Publié: (2025)
par: Kermani, Arshia, et autres
Publié: (2025)
Efficient Reinforcement Learning for Routing Jobs in Heterogeneous Queueing Systems
par: Jali, Neharika, et autres
Publié: (2024)
par: Jali, Neharika, et autres
Publié: (2024)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
par: Zheng, Wenhao, et autres
Publié: (2025)
par: Zheng, Wenhao, et autres
Publié: (2025)
Documents similaires
-
DynaSplit: A Hardware-Software Co-Design Framework for Energy-Aware Inference on Edge
par: May, Daniel, et autres
Publié: (2024) -
GREEN-CODE: Learning to Optimize Energy Efficiency in LLM-based Code Generation
par: Ilager, Shashikant, et autres
Publié: (2025) -
A Decentralized and Self-Adaptive Approach for Monitoring Volatile Edge Environments
par: Ilager, Shashikant, et autres
Publié: (2024) -
Characterizing LLM Inference Energy-Performance Tradeoffs across Workloads and GPU Scaling
par: Maliakel, Paul Joe, et autres
Publié: (2025) -
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
par: Lei, Jianlong, et autres
Publié: (2026)