BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hu, Bodun, Li, Jiamin, Xu, Le, Lee, Myungjin, Jajoo, Akshay, Kim, Geon-Woo, Xu, Hong, Akella, Aditya |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Enabling Elastic Model Serving with MultiWorld
par: Lee, Myungjin, et autres
Publié: (2024)
par: Lee, Myungjin, et autres
Publié: (2024)
Patchwork: A Unified Framework for RAG Serving
par: Hu, Bodun, et autres
Publié: (2025)
par: Hu, Bodun, et autres
Publié: (2025)
Software-Defined Agentic Serving
par: Agarwal, Saurabh, et autres
Publié: (2026)
par: Agarwal, Saurabh, et autres
Publié: (2026)
OMEGA: A Low-Latency GNN Serving System for Large Graphs
par: Kim, Geon-Woo, et autres
Publié: (2025)
par: Kim, Geon-Woo, et autres
Publié: (2025)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
par: Chung, Jae-Won, et autres
Publié: (2026)
par: Chung, Jae-Won, et autres
Publié: (2026)
Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving
par: Ma, Jeff J., et autres
Publié: (2025)
par: Ma, Jeff J., et autres
Publié: (2025)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
par: Wang, Shaoyu, et autres
Publié: (2025)
par: Wang, Shaoyu, et autres
Publié: (2025)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
par: Wu, Jingfeng, et autres
Publié: (2025)
par: Wu, Jingfeng, et autres
Publié: (2025)
CUCo: An Agentic Framework for Compute and Communication Co-design
par: Hu, Bodun, et autres
Publié: (2026)
par: Hu, Bodun, et autres
Publié: (2026)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
par: Agarwal, Saurabh, et autres
Publié: (2024)
par: Agarwal, Saurabh, et autres
Publié: (2024)
Atomicity and Abstraction for Cross-Blockchain Interactions
par: Lu, Huaixi, et autres
Publié: (2024)
par: Lu, Huaixi, et autres
Publié: (2024)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: He, Yiyuan, et autres
Publié: (2025)
par: He, Yiyuan, et autres
Publié: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
par: Mo, Zizhao, et autres
Publié: (2025)
par: Mo, Zizhao, et autres
Publié: (2025)
Nalar: An agent serving framework
par: Laju, Marco, et autres
Publié: (2026)
par: Laju, Marco, et autres
Publié: (2026)
Cloud Native System for LLM Inference Serving
par: Xu, Minxian, et autres
Publié: (2025)
par: Xu, Minxian, et autres
Publié: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
par: Hu, Jianmin, et autres
Publié: (2025)
par: Hu, Jianmin, et autres
Publié: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
par: Hu, Junhao, et autres
Publié: (2025)
par: Hu, Junhao, et autres
Publié: (2025)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
par: Li, Cong, et autres
Publié: (2025)
par: Li, Cong, et autres
Publié: (2025)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
par: He, Yiyuan, et autres
Publié: (2024)
par: He, Yiyuan, et autres
Publié: (2024)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
par: Kim, Joon Ha, et autres
Publié: (2026)
par: Kim, Joon Ha, et autres
Publié: (2026)
LobRA: Multi-tenant Fine-tuning over Heterogeneous Data
par: Lin, Sheng, et autres
Publié: (2025)
par: Lin, Sheng, et autres
Publié: (2025)
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
par: Mo, Zizhao, et autres
Publié: (2026)
par: Mo, Zizhao, et autres
Publié: (2026)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
par: Xu, Jiale, et autres
Publié: (2025)
par: Xu, Jiale, et autres
Publié: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
par: Liao, Junhan, et autres
Publié: (2025)
par: Liao, Junhan, et autres
Publié: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
par: He, Wenhao, et autres
Publié: (2026)
par: He, Wenhao, et autres
Publié: (2026)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
par: Peng, You, et autres
Publié: (2026)
par: Peng, You, et autres
Publié: (2026)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
par: Lin, Zejia, et autres
Publié: (2025)
par: Lin, Zejia, et autres
Publié: (2025)
GENSERVE: Efficient Co-Serving of Heterogeneous Diffusion Model Workloads
par: Ye, Fanjiang, et autres
Publié: (2026)
par: Ye, Fanjiang, et autres
Publié: (2026)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
par: Zhong, Yinmin, et autres
Publié: (2024)
par: Zhong, Yinmin, et autres
Publié: (2024)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
par: Sun, Tingyang, et autres
Publié: (2026)
par: Sun, Tingyang, et autres
Publié: (2026)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
par: Ruan, Chaoyi, et autres
Publié: (2025)
par: Ruan, Chaoyi, et autres
Publié: (2025)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
par: Duan, Jiangfei, et autres
Publié: (2024)
par: Duan, Jiangfei, et autres
Publié: (2024)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
par: Shen, Haiying, et autres
Publié: (2024)
par: Shen, Haiying, et autres
Publié: (2024)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
par: Jiang, Youhe, et autres
Publié: (2025)
par: Jiang, Youhe, et autres
Publié: (2025)
Predictable LLM Serving on GPU Clusters
par: Darzi, Erfan, et autres
Publié: (2025)
par: Darzi, Erfan, et autres
Publié: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
par: Cheng, Ke, et autres
Publié: (2024)
par: Cheng, Ke, et autres
Publié: (2024)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
par: Yoon, Dongha, et autres
Publié: (2025)
par: Yoon, Dongha, et autres
Publié: (2025)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
par: Lin, Yanying, et autres
Publié: (2025)
par: Lin, Yanying, et autres
Publié: (2025)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
par: Da, Wei, et autres
Publié: (2025)
par: Da, Wei, et autres
Publié: (2025)
Documents similaires
-
Enabling Elastic Model Serving with MultiWorld
par: Lee, Myungjin, et autres
Publié: (2024) -
Patchwork: A Unified Framework for RAG Serving
par: Hu, Bodun, et autres
Publié: (2025) -
Software-Defined Agentic Serving
par: Agarwal, Saurabh, et autres
Publié: (2026) -
OMEGA: A Low-Latency GNN Serving System for Large Graphs
par: Kim, Geon-Woo, et autres
Publié: (2025) -
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
par: Chung, Jae-Won, et autres
Publié: (2026)