EcoServe: Designing Carbon-Aware AI Inference Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yueying, Hu, Zhanqiu, Choukse, Esha, Fonseca, Rodrigo, Suh, G. Edward, Gupta, Udit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
von: Du, Jiangsu, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Compound AI Systems
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
von: Qiu, Haoran, et al.
Veröffentlicht: (2026)
von: Qiu, Haoran, et al.
Veröffentlicht: (2026)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
von: Saurez, Enrique, et al.
Veröffentlicht: (2024)
von: Saurez, Enrique, et al.
Veröffentlicht: (2024)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
Aladdin: Joint Placement and Scaling for SLO-Aware LLM Serving
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
von: Nie, Chengyi, et al.
Veröffentlicht: (2024)
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
von: Oviedo, Felipe, et al.
Veröffentlicht: (2025)
von: Oviedo, Felipe, et al.
Veröffentlicht: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
von: Hu, Jianmin, et al.
Veröffentlicht: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
von: Nguyen, Thanh-Tung, et al.
Veröffentlicht: (2025)
von: Nguyen, Thanh-Tung, et al.
Veröffentlicht: (2025)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
EcoLife: Carbon-Aware Serverless Function Scheduling for Sustainable Computing
von: Jiang, Yankai, et al.
Veröffentlicht: (2024)
von: Jiang, Yankai, et al.
Veröffentlicht: (2024)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
von: Wolfrath, Joel, et al.
Veröffentlicht: (2025)
von: Wolfrath, Joel, et al.
Veröffentlicht: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
von: Du, Boxiao, et al.
Veröffentlicht: (2026)
Taming the Memory Footprint Crisis: System Design for Production Diffusion LLM Serving
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
von: Zhan, Huiyou, et al.
Veröffentlicht: (2025)
von: Zhan, Huiyou, et al.
Veröffentlicht: (2025)
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
von: Razavi, Kamran, et al.
Veröffentlicht: (2024)
EcoShift: Performance-Aware Power Management for Power-Constrained Heterogeneous Systems
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
von: He, Yiyuan, et al.
Veröffentlicht: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
Shelby: Decentralized Storage Designed to Serve
von: Goren, Guy, et al.
Veröffentlicht: (2025)
von: Goren, Guy, et al.
Veröffentlicht: (2025)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
von: Ahmad, Sohaib, et al.
Veröffentlicht: (2024)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
von: Xia, Yifei, et al.
Veröffentlicht: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
Intent-based System Design and Operation
von: Anand, Vaastav, et al.
Veröffentlicht: (2025)
von: Anand, Vaastav, et al.
Veröffentlicht: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
von: He, Yiyuan, et al.
Veröffentlicht: (2024)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
von: Basit, Omar, et al.
Veröffentlicht: (2026)
von: Basit, Omar, et al.
Veröffentlicht: (2026)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
Kairos: A Scalable Serving System for Physical AI
von: Dai, Yinwei, et al.
Veröffentlicht: (2026)
von: Dai, Yinwei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
von: Du, Jiangsu, et al.
Veröffentlicht: (2025) -
Towards Resource-Efficient Compound AI Systems
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025) -
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
von: Qiu, Haoran, et al.
Veröffentlicht: (2026) -
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025) -
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)