Towards Resource-Efficient Compound AI Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Chaudhry, Gohar Irfan, Choukse, Esha, Goiri, Íñigo, Fonseca, Rodrigo, Belay, Adam, Bianchini, Ricardo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
di: Qiu, Haoran, et al.
Pubblicazione: (2026)
di: Qiu, Haoran, et al.
Pubblicazione: (2026)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
di: Saurez, Enrique, et al.
Pubblicazione: (2024)
di: Saurez, Enrique, et al.
Pubblicazione: (2024)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
di: Qiu, Haoran, et al.
Pubblicazione: (2025)
di: Qiu, Haoran, et al.
Pubblicazione: (2025)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
Benchmarking Compound AI Applications for Hardware-Software Co-Design
di: Samuthrsindh, Paramuth, et al.
Pubblicazione: (2026)
di: Samuthrsindh, Paramuth, et al.
Pubblicazione: (2026)
EcoServe: Designing Carbon-Aware AI Inference Systems
di: Li, Yueying, et al.
Pubblicazione: (2025)
di: Li, Yueying, et al.
Pubblicazione: (2025)
Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
Cloud abstractions for AI workloads
di: Canini, Marco, et al.
Pubblicazione: (2025)
di: Canini, Marco, et al.
Pubblicazione: (2025)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
di: Agrawal, Amey, et al.
Pubblicazione: (2024)
di: Agrawal, Amey, et al.
Pubblicazione: (2024)
Designing Datacenter Power Delivery Hierarchies for the AI Era
di: Wilkins, Grant, et al.
Pubblicazione: (2026)
di: Wilkins, Grant, et al.
Pubblicazione: (2026)
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
di: Oviedo, Felipe, et al.
Pubblicazione: (2025)
di: Oviedo, Felipe, et al.
Pubblicazione: (2025)
The AI_INFN Platform: Artificial Intelligence Development in the Cloud
di: Anderlini, Lucio, et al.
Pubblicazione: (2025)
di: Anderlini, Lucio, et al.
Pubblicazione: (2025)
Towards Cloud Efficiency with Large-scale Workload Characterization
di: Parayil, Anjaly, et al.
Pubblicazione: (2024)
di: Parayil, Anjaly, et al.
Pubblicazione: (2024)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
di: Jain, Kunal, et al.
Pubblicazione: (2024)
di: Jain, Kunal, et al.
Pubblicazione: (2024)
AI-Driven Cloud Resource Optimization for Multi-Cluster Environments
di: Punniyamoorthy, Vinoth, et al.
Pubblicazione: (2025)
di: Punniyamoorthy, Vinoth, et al.
Pubblicazione: (2025)
Adaptive AI-based Decentralized Resource Management in the Cloud-Edge Continuum
di: Li, Lanpei, et al.
Pubblicazione: (2025)
di: Li, Lanpei, et al.
Pubblicazione: (2025)
Percepta: High Performance Stream Processing at the Edge
di: Sousa, Clarisse, et al.
Pubblicazione: (2025)
di: Sousa, Clarisse, et al.
Pubblicazione: (2025)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
di: Huang, Lexiang, et al.
Pubblicazione: (2024)
di: Huang, Lexiang, et al.
Pubblicazione: (2024)
KAITIAN: A Unified Communication Framework for Enabling Efficient Collaboration Across Heterogeneous Accelerators in Embodied AI Systems
di: Lin, Jieke, et al.
Pubblicazione: (2025)
di: Lin, Jieke, et al.
Pubblicazione: (2025)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
di: Xie, Jincheng, et al.
Pubblicazione: (2026)
The (R)evolution of Scientific Workflows in the Agentic AI Era: Towards Autonomous Science
di: Shin, Woong, et al.
Pubblicazione: (2025)
di: Shin, Woong, et al.
Pubblicazione: (2025)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
di: Yuan, Yichao, et al.
Pubblicazione: (2025)
di: Yuan, Yichao, et al.
Pubblicazione: (2025)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
di: Liu, Wentao, et al.
Pubblicazione: (2025)
di: Liu, Wentao, et al.
Pubblicazione: (2025)
Towards using Reinforcement Learning for Scaling and Data Replication in Cloud Systems
di: Mokadem, Riad, et al.
Pubblicazione: (2024)
di: Mokadem, Riad, et al.
Pubblicazione: (2024)
An AI-Driven Framework for Energy-Efficient Environmental Monitoring in Smart Cities Using Edge Intelligence
di: Liu, Yichen, et al.
Pubblicazione: (2026)
di: Liu, Yichen, et al.
Pubblicazione: (2026)
Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems
di: Alshaer, Samer, et al.
Pubblicazione: (2025)
di: Alshaer, Samer, et al.
Pubblicazione: (2025)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
di: Kamahori, Keisuke, et al.
Pubblicazione: (2026)
di: Kamahori, Keisuke, et al.
Pubblicazione: (2026)
Joint Resource Optimization, Computation Offloading and Resource Slicing for Multi-Edge Traffic-Cognitive Networks
di: Xiaoyang, Ting, et al.
Pubblicazione: (2024)
di: Xiaoyang, Ting, et al.
Pubblicazione: (2024)
Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
di: Fedorov, Igor, et al.
Pubblicazione: (2024)
di: Fedorov, Igor, et al.
Pubblicazione: (2024)
Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and Certification
di: Ranjbar, Behnaz, et al.
Pubblicazione: (2026)
di: Ranjbar, Behnaz, et al.
Pubblicazione: (2026)
Compass: Optimizing Compound AI Workflows for Dynamic Adaptation
di: Gravara, Milos, et al.
Pubblicazione: (2026)
di: Gravara, Milos, et al.
Pubblicazione: (2026)
Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
di: Ye, Shengyuan, et al.
Pubblicazione: (2025)
di: Ye, Shengyuan, et al.
Pubblicazione: (2025)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
di: Wu, Qi, et al.
Pubblicazione: (2026)
di: Wu, Qi, et al.
Pubblicazione: (2026)
Hardware-Aware Reformulation of Convolutions for Efficient Execution on Specialized AI Hardware: A Case Study on NVIDIA Tensor Cores
di: Bikshandi, Ganesh
Pubblicazione: (2026)
di: Bikshandi, Ganesh
Pubblicazione: (2026)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
di: Yu, Dianhai, et al.
Pubblicazione: (2022)
di: Yu, Dianhai, et al.
Pubblicazione: (2022)
Efficient and Scalable Agentic AI with Heterogeneous Systems
di: Asgar, Zain, et al.
Pubblicazione: (2025)
di: Asgar, Zain, et al.
Pubblicazione: (2025)
Documenti analoghi
-
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
di: Qiu, Haoran, et al.
Pubblicazione: (2026) -
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
di: Saurez, Enrique, et al.
Pubblicazione: (2024) -
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025) -
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024) -
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
di: Qiu, Haoran, et al.
Pubblicazione: (2025)