Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhao, Lingxiao, Zhou, Haoran, Che, Yuezhi, Cheng, Dazhao |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
RAPID-Serve: Resource-efficient and Accelerated P/D Intra-GPU Disaggregation
par: Masood, Amna, et autres
Publié: (2026)
par: Masood, Amna, et autres
Publié: (2026)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
par: Woo, Sunghyeon, et autres
Publié: (2026)
par: Woo, Sunghyeon, et autres
Publié: (2026)
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
par: Liang, Yunkai, et autres
Publié: (2025)
par: Liang, Yunkai, et autres
Publié: (2025)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
par: Shi, Xiaoxiang, et autres
Publié: (2025)
par: Shi, Xiaoxiang, et autres
Publié: (2025)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
par: Wang, Chong, et autres
Publié: (2026)
par: Wang, Chong, et autres
Publié: (2026)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
par: Zhang, Zeyu, et autres
Publié: (2025)
par: Zhang, Zeyu, et autres
Publié: (2025)
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
par: Oh, Hyungjun, et autres
Publié: (2024)
par: Oh, Hyungjun, et autres
Publié: (2024)
Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
par: Luo, Yizhou, et autres
Publié: (2024)
par: Luo, Yizhou, et autres
Publié: (2024)
KVDirect: Distributed Disaggregated LLM Inference
par: Chen, Shiyang, et autres
Publié: (2024)
par: Chen, Shiyang, et autres
Publié: (2024)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
par: Zhong, Shuzhang, et autres
Publié: (2025)
par: Zhong, Shuzhang, et autres
Publié: (2025)
Timeliness-Oriented Scheduling and Resource Allocation in Multi-Region Collaborative Perception
par: Zhu, Mengmeng, et autres
Publié: (2026)
par: Zhu, Mengmeng, et autres
Publié: (2026)
FIKIT: Priority-Based Real-time GPU Multi-tasking Scheduling with Kernel Identification
par: Wu, Wenqing
Publié: (2023)
par: Wu, Wenqing
Publié: (2023)
A Practical Two-Stage Framework for GPU Resource and Power Prediction in Heterogeneous HPC Systems
par: Oztop, Beste, et autres
Publié: (2026)
par: Oztop, Beste, et autres
Publié: (2026)
SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
par: Li, Luchang, et autres
Publié: (2026)
par: Li, Luchang, et autres
Publié: (2026)
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
par: Wang, Zerui, et autres
Publié: (2025)
par: Wang, Zerui, et autres
Publié: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
par: Zhu, Ruidong, et autres
Publié: (2025)
par: Zhu, Ruidong, et autres
Publié: (2025)
WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM Serving
par: Lou, Chiheng, et autres
Publié: (2025)
par: Lou, Chiheng, et autres
Publié: (2025)
MultiTASC++: A Continuously Adaptive Scheduler for Edge-Based Multi-Device Cascade Inference
par: Nikolaidis, Sokratis, et autres
Publié: (2024)
par: Nikolaidis, Sokratis, et autres
Publié: (2024)
ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache
par: Wang, Shao, et autres
Publié: (2026)
par: Wang, Shao, et autres
Publié: (2026)
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
par: Ding, Jianru, et autres
Publié: (2026)
par: Ding, Jianru, et autres
Publié: (2026)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
par: Luo, Ziyue, et autres
Publié: (2025)
par: Luo, Ziyue, et autres
Publié: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
par: Yu, Minchen, et autres
Publié: (2023)
par: Yu, Minchen, et autres
Publié: (2023)
StraightLine: An End-to-End Resource-Aware Scheduler for Machine Learning Application Requests
par: Ching, Cheng-Wei, et autres
Publié: (2024)
par: Ching, Cheng-Wei, et autres
Publié: (2024)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
par: Wu, Yu, et autres
Publié: (2025)
par: Wu, Yu, et autres
Publié: (2025)
GPU Cluster Scheduling for Network-Sensitive Deep Learning
par: Sharma, Aakash, et autres
Publié: (2024)
par: Sharma, Aakash, et autres
Publié: (2024)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
par: Zhong, Yinmin, et autres
Publié: (2025)
par: Zhong, Yinmin, et autres
Publié: (2025)
DHP: Efficient Scaling of MLLM Training with Dynamic Hybrid Parallelism
par: Niu, Yifan, et autres
Publié: (2026)
par: Niu, Yifan, et autres
Publié: (2026)
Optimal Scheduling Algorithms for LLM Inference: Theory and Practice
par: Bari, Agrim, et autres
Publié: (2025)
par: Bari, Agrim, et autres
Publié: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
par: Zhang, Haolin, et autres
Publié: (2025)
par: Zhang, Haolin, et autres
Publié: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
par: Zhao, Bohan, et autres
Publié: (2025)
par: Zhao, Bohan, et autres
Publié: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
par: Wang, Qipeng
Publié: (2026)
par: Wang, Qipeng
Publié: (2026)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
par: Zhang, Zeyu, et autres
Publié: (2024)
par: Zhang, Zeyu, et autres
Publié: (2024)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
par: Recasens, Pol G., et autres
Publié: (2025)
par: Recasens, Pol G., et autres
Publié: (2025)
GPU-Accelerated Optimization of Transformer-Based Neural Networks for Real-Time Inference
par: Mukherjee, Soutrik, et autres
Publié: (2026)
par: Mukherjee, Soutrik, et autres
Publié: (2026)
Reinforcement Learning for Adaptive Resource Scheduling in Complex System Environments
par: Li, Pochun, et autres
Publié: (2024)
par: Li, Pochun, et autres
Publié: (2024)
SAIR: Cost-Efficient Multi-Stage ML Pipeline Autoscaling via In-Context Reinforcement Learning
par: Su, Jianchang, et autres
Publié: (2026)
par: Su, Jianchang, et autres
Publié: (2026)
SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
par: Zhang, Hengrui, et autres
Publié: (2025)
par: Zhang, Hengrui, et autres
Publié: (2025)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
par: Hu, Tiancheng, et autres
Publié: (2026)
par: Hu, Tiancheng, et autres
Publié: (2026)
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
par: Zhang, Tianyi, et autres
Publié: (2025)
par: Zhang, Tianyi, et autres
Publié: (2025)
Metadata-Guided Adaptable Frequency Scaling across Heterogeneous Applications and Devices
par: Yan, Jinqi, et autres
Publié: (2025)
par: Yan, Jinqi, et autres
Publié: (2025)
Documents similaires
-
RAPID-Serve: Resource-efficient and Accelerated P/D Intra-GPU Disaggregation
par: Masood, Amna, et autres
Publié: (2026) -
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
par: Woo, Sunghyeon, et autres
Publié: (2026) -
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
par: Liang, Yunkai, et autres
Publié: (2025) -
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
par: Shi, Xiaoxiang, et autres
Publié: (2025) -
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
par: Wang, Chong, et autres
Publié: (2026)