Data Driven Optimization of GPU efficiency for Distributed LLM Adapter Serving
Fuente:
arXiv
Guardado en:
| Autores principales: | Agullo, Ferran, Oliveras, Joan, Wang, Chen, Gutierrez-Torre, Alberto, Tardieu, Olivier, Youssef, Alaa, Torres, Jordi, Berral, Josep Ll. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
por: Recasens, Pol G., et al.
Publicado: (2025)
por: Recasens, Pol G., et al.
Publicado: (2025)
Enabling an OpenStack-based cloud on top of RISC-V hardware
por: Marrón, Diego, et al.
Publicado: (2024)
por: Marrón, Diego, et al.
Publicado: (2024)
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
por: Agullo, Ferran, et al.
Publicado: (2025)
por: Agullo, Ferran, et al.
Publicado: (2025)
Predictable LLM Serving on GPU Clusters
por: Darzi, Erfan, et al.
Publicado: (2025)
por: Darzi, Erfan, et al.
Publicado: (2025)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
por: Jiang, Youhe, et al.
Publicado: (2025)
por: Jiang, Youhe, et al.
Publicado: (2025)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
por: Lin, Zejia, et al.
Publicado: (2025)
por: Lin, Zejia, et al.
Publicado: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
por: Mo, Zizhao, et al.
Publicado: (2026)
por: Mo, Zizhao, et al.
Publicado: (2026)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
por: Bai, Fengyao, et al.
Publicado: (2026)
por: Bai, Fengyao, et al.
Publicado: (2026)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
por: Zhao, Bohan, et al.
Publicado: (2025)
por: Zhao, Bohan, et al.
Publicado: (2025)
ExpertWeave: Efficiently Serving Expert-Specialized Fine-Tuned Adapters at Scale
por: Shi, Ge, et al.
Publicado: (2025)
por: Shi, Ge, et al.
Publicado: (2025)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
por: Lou, Chiheng, et al.
Publicado: (2025)
por: Lou, Chiheng, et al.
Publicado: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
por: Zhang, Yuning, et al.
Publicado: (2026)
por: Zhang, Yuning, et al.
Publicado: (2026)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
por: Qiao, Yifan, et al.
Publicado: (2024)
por: Qiao, Yifan, et al.
Publicado: (2024)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
por: Mo, Zizhao, et al.
Publicado: (2025)
por: Mo, Zizhao, et al.
Publicado: (2025)
RAPID-Serve: Resource-efficient and Accelerated P/D Intra-GPU Disaggregation
por: Masood, Amna, et al.
Publicado: (2026)
por: Masood, Amna, et al.
Publicado: (2026)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
por: Ruan, Chaoyi, et al.
Publicado: (2025)
por: Ruan, Chaoyi, et al.
Publicado: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
por: Duan, Jiangfei, et al.
Publicado: (2024)
por: Duan, Jiangfei, et al.
Publicado: (2024)
Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
por: Lv, Cunchi, et al.
Publicado: (2025)
por: Lv, Cunchi, et al.
Publicado: (2025)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
por: Zhan, Huiyou, et al.
Publicado: (2025)
por: Zhan, Huiyou, et al.
Publicado: (2025)
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
por: Scheinert, Dominik, et al.
Publicado: (2026)
por: Scheinert, Dominik, et al.
Publicado: (2026)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
por: Hu, Cunchen, et al.
Publicado: (2024)
por: Hu, Cunchen, et al.
Publicado: (2024)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
por: Shen, Haiying, et al.
Publicado: (2024)
por: Shen, Haiying, et al.
Publicado: (2024)
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
por: Agrawal, Amey, et al.
Publicado: (2026)
por: Agrawal, Amey, et al.
Publicado: (2026)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
por: Shi, Xiaoxiang, et al.
Publicado: (2025)
por: Shi, Xiaoxiang, et al.
Publicado: (2025)
HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
por: Lou, Chiheng, et al.
Publicado: (2025)
por: Lou, Chiheng, et al.
Publicado: (2025)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
por: Zhou, Qihui, et al.
Publicado: (2025)
por: Zhou, Qihui, et al.
Publicado: (2025)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
por: Jeon, Beomyeol, et al.
Publicado: (2024)
por: Jeon, Beomyeol, et al.
Publicado: (2024)
Cloud Native System for LLM Inference Serving
por: Xu, Minxian, et al.
Publicado: (2025)
por: Xu, Minxian, et al.
Publicado: (2025)
Fast State Restoration in LLM Serving with HCache
por: Gao, Shiwei, et al.
Publicado: (2024)
por: Gao, Shiwei, et al.
Publicado: (2024)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
por: Li, Suyi, et al.
Publicado: (2024)
por: Li, Suyi, et al.
Publicado: (2024)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
por: Du, Boxiao, et al.
Publicado: (2026)
por: Du, Boxiao, et al.
Publicado: (2026)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
por: Du, Jiangsu, et al.
Publicado: (2025)
por: Du, Jiangsu, et al.
Publicado: (2025)
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
por: Kong, Z. Jonny, et al.
Publicado: (2025)
por: Kong, Z. Jonny, et al.
Publicado: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
por: Qianli, Liu, et al.
Publicado: (2025)
por: Qianli, Liu, et al.
Publicado: (2025)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
por: Xu, Jiale, et al.
Publicado: (2025)
por: Xu, Jiale, et al.
Publicado: (2025)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
por: Zhang, Chen, et al.
Publicado: (2025)
por: Zhang, Chen, et al.
Publicado: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
por: He, Yiyuan, et al.
Publicado: (2025)
por: He, Yiyuan, et al.
Publicado: (2025)
An efficient GPU approach for designing 3D cultural heritage information systems
por: López, Luis, et al.
Publicado: (2025)
por: López, Luis, et al.
Publicado: (2025)
Ejemplares similares
-
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
por: Recasens, Pol G., et al.
Publicado: (2025) -
Enabling an OpenStack-based cloud on top of RISC-V hardware
por: Marrón, Diego, et al.
Publicado: (2024) -
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
por: Agullo, Ferran, et al.
Publicado: (2025) -
Predictable LLM Serving on GPU Clusters
por: Darzi, Erfan, et al.
Publicado: (2025) -
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
por: Jiang, Youhe, et al.
Publicado: (2025)