ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
Fuente:
arXiv
Salvato in:
| Autori principali: | Qiu, Haoran, Biswas, Anish, Zhao, Zihan, Mohan, Jayashree, Khare, Alind, Choukse, Esha, Goiri, Íñigo, Zhang, Zeyu, Shen, Haiying, Bansal, Chetan, Ramjee, Ramachandran, Fonseca, Rodrigo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
di: Biswas, Anish, et al.
Pubblicazione: (2026)
di: Biswas, Anish, et al.
Pubblicazione: (2026)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
di: Qiu, Haoran, et al.
Pubblicazione: (2026)
di: Qiu, Haoran, et al.
Pubblicazione: (2026)
Niyama : Breaking the Silos of LLM Inference Serving
di: Goel, Kanishk, et al.
Pubblicazione: (2025)
di: Goel, Kanishk, et al.
Pubblicazione: (2025)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
di: Prabhu, Ramya, et al.
Pubblicazione: (2024)
di: Prabhu, Ramya, et al.
Pubblicazione: (2024)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
di: Agrawal, Amey, et al.
Pubblicazione: (2024)
di: Agrawal, Amey, et al.
Pubblicazione: (2024)
On Evaluating Performance of LLM Inference Serving Systems
di: Agrawal, Amey, et al.
Pubblicazione: (2025)
di: Agrawal, Amey, et al.
Pubblicazione: (2025)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
di: Agrawal, Amey, et al.
Pubblicazione: (2024)
di: Agrawal, Amey, et al.
Pubblicazione: (2024)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
Towards Resource-Efficient Compound AI Systems
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
EcoServe: Designing Carbon-Aware AI Inference Systems
di: Li, Yueying, et al.
Pubblicazione: (2025)
di: Li, Yueying, et al.
Pubblicazione: (2025)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
Sherlock: Reliable and Efficient Agentic Workflow Execution
di: Ro, Yeonju, et al.
Pubblicazione: (2025)
di: Ro, Yeonju, et al.
Pubblicazione: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
di: Patel, Pratyush, et al.
Pubblicazione: (2023)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
di: Saurez, Enrique, et al.
Pubblicazione: (2024)
di: Saurez, Enrique, et al.
Pubblicazione: (2024)
ASTRA: Accurate and Scalable ANNS-based Training of Extreme Classifiers
di: Mehta, Sonu, et al.
Pubblicazione: (2024)
di: Mehta, Sonu, et al.
Pubblicazione: (2024)
AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
di: Barke, Shraddha, et al.
Pubblicazione: (2026)
di: Barke, Shraddha, et al.
Pubblicazione: (2026)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
di: Shen, Haiying, et al.
Pubblicazione: (2024)
di: Shen, Haiying, et al.
Pubblicazione: (2024)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
di: Jain, Kunal, et al.
Pubblicazione: (2024)
di: Jain, Kunal, et al.
Pubblicazione: (2024)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
di: Hu, Xiannan, et al.
Pubblicazione: (2025)
di: Hu, Xiannan, et al.
Pubblicazione: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
di: Jin, Yibo, et al.
Pubblicazione: (2024)
di: Jin, Yibo, et al.
Pubblicazione: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
di: Hu, Cunchen, et al.
Pubblicazione: (2024)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
di: Kamath, Aditya K, et al.
Pubblicazione: (2024)
di: Kamath, Aditya K, et al.
Pubblicazione: (2024)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
di: Bai, Fan, et al.
Pubblicazione: (2026)
di: Bai, Fan, et al.
Pubblicazione: (2026)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
di: Su, Zhaoyuan, et al.
Pubblicazione: (2025)
Trinity: Disaggregating Vector Search from Prefill-Decode Disaggregation in LLM Serving
di: Liu, Yi, et al.
Pubblicazione: (2025)
di: Liu, Yi, et al.
Pubblicazione: (2025)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
Towards Cloud Efficiency with Large-scale Workload Characterization
di: Parayil, Anjaly, et al.
Pubblicazione: (2024)
di: Parayil, Anjaly, et al.
Pubblicazione: (2024)
Input-Dependent Power Usage in GPUs
di: Gregersen, Theo, et al.
Pubblicazione: (2024)
di: Gregersen, Theo, et al.
Pubblicazione: (2024)
DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
di: Liu, Yuhan, et al.
Pubblicazione: (2024)
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
di: Qin, Ruoyu, et al.
Pubblicazione: (2024)
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
di: Ding, Jianru, et al.
Pubblicazione: (2026)
di: Ding, Jianru, et al.
Pubblicazione: (2026)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
Efficiently Serving Large Multimodal Models Using EPD Disaggregation
di: Singh, Gursimran, et al.
Pubblicazione: (2024)
di: Singh, Gursimran, et al.
Pubblicazione: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
di: He, Yiyuan, et al.
Pubblicazione: (2025)
di: He, Yiyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
di: Biswas, Anish, et al.
Pubblicazione: (2026) -
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
di: Qiu, Haoran, et al.
Pubblicazione: (2026) -
Niyama : Breaking the Silos of LLM Inference Serving
di: Goel, Kanishk, et al.
Pubblicazione: (2025) -
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
di: Prabhu, Ramya, et al.
Pubblicazione: (2024) -
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
di: Agrawal, Amey, et al.
Pubblicazione: (2024)