No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Agrawal, Amey, Qiu, Haoran, Chen, Junda, Goiri, Íñigo, Zhang, Chaojie, Shahid, Rayyan, Ramjee, Ramachandran, Tumanov, Alexey, Choukse, Esha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
von: Qiu, Haoran, et al.
Veröffentlicht: (2026)
von: Qiu, Haoran, et al.
Veröffentlicht: (2026)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
On Evaluating Performance of LLM Inference Serving Systems
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
Splitwise: Efficient generative LLM inference using phase splitting
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
von: Patel, Pratyush, et al.
Veröffentlicht: (2023)
Towards Resource-Efficient Compound AI Systems
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
von: Saurez, Enrique, et al.
Veröffentlicht: (2024)
von: Saurez, Enrique, et al.
Veröffentlicht: (2024)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference
von: Deshmukh, Dhruv, et al.
Veröffentlicht: (2025)
von: Deshmukh, Dhruv, et al.
Veröffentlicht: (2025)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025)
von: Gond, Raja, et al.
Veröffentlicht: (2025)
EcoServe: Designing Carbon-Aware AI Inference Systems
von: Li, Yueying, et al.
Veröffentlicht: (2025)
von: Li, Yueying, et al.
Veröffentlicht: (2025)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
von: Gond, Raja, et al.
Veröffentlicht: (2026)
von: Gond, Raja, et al.
Veröffentlicht: (2026)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
von: Jain, Kunal, et al.
Veröffentlicht: (2024)
von: Jain, Kunal, et al.
Veröffentlicht: (2024)
SuperFedNAS: Cost-Efficient Federated Neural Architecture Search for On-Device Inference
von: Khare, Alind, et al.
Veröffentlicht: (2023)
von: Khare, Alind, et al.
Veröffentlicht: (2023)
Vidur: A Large-Scale Simulation Framework For LLM Inference
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
Niyama : Breaking the Silos of LLM Inference Serving
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
von: Yarlagadda, Srihas, et al.
Veröffentlicht: (2025)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
von: Biswas, Anish, et al.
Veröffentlicht: (2026)
von: Biswas, Anish, et al.
Veröffentlicht: (2026)
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
von: Oviedo, Felipe, et al.
Veröffentlicht: (2025)
von: Oviedo, Felipe, et al.
Veröffentlicht: (2025)
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
von: Agrawal, Amey, et al.
Veröffentlicht: (2026)
von: Agrawal, Amey, et al.
Veröffentlicht: (2026)
Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
Towards Cloud Efficiency with Large-scale Workload Characterization
von: Parayil, Anjaly, et al.
Veröffentlicht: (2024)
von: Parayil, Anjaly, et al.
Veröffentlicht: (2024)
Cloud abstractions for AI workloads
von: Canini, Marco, et al.
Veröffentlicht: (2025)
von: Canini, Marco, et al.
Veröffentlicht: (2025)
Tackling Privacy Heterogeneity in Differentially Private Federated Learning
von: Xu, Ruichen, et al.
Veröffentlicht: (2026)
von: Xu, Ruichen, et al.
Veröffentlicht: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
von: Wang, Qipeng
Veröffentlicht: (2026)
von: Wang, Qipeng
Veröffentlicht: (2026)
PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping
von: Zhao, Zhixin, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixin, et al.
Veröffentlicht: (2026)
Sherlock: Reliable and Efficient Agentic Workflow Execution
von: Ro, Yeonju, et al.
Veröffentlicht: (2025)
von: Ro, Yeonju, et al.
Veröffentlicht: (2025)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
von: Yu, Shibo, et al.
Veröffentlicht: (2025)
von: Yu, Shibo, et al.
Veröffentlicht: (2025)
ParallelSFL: A Novel Split Federated Learning Framework Tackling Heterogeneity Issues
von: Liao, Yunming, et al.
Veröffentlicht: (2024)
von: Liao, Yunming, et al.
Veröffentlicht: (2024)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
von: Chen, Jiefei, et al.
Veröffentlicht: (2026)
von: Chen, Jiefei, et al.
Veröffentlicht: (2026)
Tackling Intertwined Data and Device Heterogeneities in Federated Learning with Unlimited Staleness
von: Wang, Haoming, et al.
Veröffentlicht: (2023)
von: Wang, Haoming, et al.
Veröffentlicht: (2023)
Nanvix: A Multikernel OS Design for High-Density Serverless Deployments
von: Segarra, Carlos, et al.
Veröffentlicht: (2026)
von: Segarra, Carlos, et al.
Veröffentlicht: (2026)
Tackling Resource-Constrained and Data-Heterogeneity in Federated Learning with Double-Weight Sparse Pack
von: Yang, Qiantao, et al.
Veröffentlicht: (2026)
von: Yang, Qiantao, et al.
Veröffentlicht: (2026)
Scaling Up Throughput-oriented LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
von: Phung, Thanh Son, et al.
Veröffentlicht: (2025)
von: Phung, Thanh Son, et al.
Veröffentlicht: (2025)
Efficiently Executing High-throughput Lightweight LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
von: Phung, Thanh Son, et al.
Veröffentlicht: (2025)
von: Phung, Thanh Son, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024) -
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024) -
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025) -
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
von: Qiu, Haoran, et al.
Veröffentlicht: (2026) -
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)