Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mo, Zizhao, Liao, Jianxiong, Xu, Huanle, Zhou, Zhi, Xu, Chengzhong |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
par: Mo, Zizhao, et autres
Publié: (2026)
par: Mo, Zizhao, et autres
Publié: (2026)
Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
par: Mo, Zizhao, et autres
Publié: (2024)
par: Mo, Zizhao, et autres
Publié: (2024)
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
par: Wu, Jingfeng, et autres
Publié: (2025)
par: Wu, Jingfeng, et autres
Publié: (2025)
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
par: Kong, Z. Jonny, et autres
Publié: (2025)
par: Kong, Z. Jonny, et autres
Publié: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
par: Hu, Jianmin, et autres
Publié: (2025)
par: Hu, Jianmin, et autres
Publié: (2025)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
par: Lin, Yanying, et autres
Publié: (2025)
par: Lin, Yanying, et autres
Publié: (2025)
Cloud Native System for LLM Inference Serving
par: Xu, Minxian, et autres
Publié: (2025)
par: Xu, Minxian, et autres
Publié: (2025)
DOPD: A Dynamic PD-Disaggregation Architecture for Maximizing Goodput in LLM Inference Serving
par: Liao, Junhan, et autres
Publié: (2025)
par: Liao, Junhan, et autres
Publié: (2025)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
par: Liang, Antian, et autres
Publié: (2025)
par: Liang, Antian, et autres
Publié: (2025)
Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
par: Li, Cong, et autres
Publié: (2025)
par: Li, Cong, et autres
Publié: (2025)
Predictable LLM Serving on GPU Clusters
par: Darzi, Erfan, et autres
Publié: (2025)
par: Darzi, Erfan, et autres
Publié: (2025)
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
par: He, Yiyuan, et autres
Publié: (2025)
par: He, Yiyuan, et autres
Publié: (2025)
Heterogeneous Federated Fine-Tuning with Parallel One-Rank Adaptation
par: Zhang, Zikai, et autres
Publié: (2026)
par: Zhang, Zikai, et autres
Publié: (2026)
UELLM: A Unified and Efficient Approach for LLM Inference Serving
par: He, Yiyuan, et autres
Publié: (2024)
par: He, Yiyuan, et autres
Publié: (2024)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
par: Zhou, Qihui, et autres
Publié: (2025)
par: Zhou, Qihui, et autres
Publié: (2025)
Cronus: Efficient LLM inference on Heterogeneous GPU Clusters via Partially Disaggregated Prefill
par: Liu, Yunzhao, et autres
Publié: (2025)
par: Liu, Yunzhao, et autres
Publié: (2025)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
par: Qiao, Yifan, et autres
Publié: (2024)
par: Qiao, Yifan, et autres
Publié: (2024)
ParallelSFL: A Novel Split Federated Learning Framework Tackling Heterogeneity Issues
par: Liao, Yunming, et autres
Publié: (2024)
par: Liao, Yunming, et autres
Publié: (2024)
DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-based Clusters
par: Bai, Haoyu, et autres
Publié: (2024)
par: Bai, Haoyu, et autres
Publié: (2024)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
par: Lin, Zejia, et autres
Publié: (2025)
par: Lin, Zejia, et autres
Publié: (2025)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
par: Xu, Minxian, et autres
Publié: (2026)
par: Xu, Minxian, et autres
Publié: (2026)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
par: Guo, Runsheng Benson, et autres
Publié: (2025)
par: Guo, Runsheng Benson, et autres
Publié: (2025)
Cephalo: Harnessing Heterogeneous GPU Clusters for Training Transformer Models
par: Guo, Runsheng Benson, et autres
Publié: (2024)
par: Guo, Runsheng Benson, et autres
Publié: (2024)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
par: Peng, You, et autres
Publié: (2026)
par: Peng, You, et autres
Publié: (2026)
Exploring Fine-grained Task Parallelism on Simultaneous Multithreading Cores
par: Los, Denis, et autres
Publié: (2024)
par: Los, Denis, et autres
Publié: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
par: Gu, Jianfeng, et autres
Publié: (2025)
par: Gu, Jianfeng, et autres
Publié: (2025)
Taming GPU Underutilization via Static Partitioning and Fine-grained CPU Offloading
par: Schieffer, Gabin, et autres
Publié: (2026)
par: Schieffer, Gabin, et autres
Publié: (2026)
BlockLLM: Multi-tenant Finer-grained Serving for Large Language Models
par: Hu, Bodun, et autres
Publié: (2024)
par: Hu, Bodun, et autres
Publié: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
par: Zhang, WenZheng, et autres
Publié: (2024)
par: Zhang, WenZheng, et autres
Publié: (2024)
Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
par: Wu, Yebo, et autres
Publié: (2024)
par: Wu, Yebo, et autres
Publié: (2024)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
par: Zhang, Chen, et autres
Publié: (2025)
par: Zhang, Chen, et autres
Publié: (2025)
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
par: Zhang, Shiwei, et autres
Publié: (2024)
par: Zhang, Shiwei, et autres
Publié: (2024)
LoHan: Low-Cost High-Performance Framework to Fine-Tune 100B Model on a Consumer GPU
par: Liao, Changyue, et autres
Publié: (2024)
par: Liao, Changyue, et autres
Publié: (2024)
APEX: An Extensible and Dynamism-Aware Simulator for Automated Parallel Execution in LLM Serving
par: Lin, Yi-Chien, et autres
Publié: (2024)
par: Lin, Yi-Chien, et autres
Publié: (2024)
Breaking the Memory Wall for Heterogeneous Federated Learning via Progressive Training
par: Wu, Yebo, et autres
Publié: (2024)
par: Wu, Yebo, et autres
Publié: (2024)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
par: Chang, Zihan, et autres
Publié: (2024)
par: Chang, Zihan, et autres
Publié: (2024)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
par: Wang, Zhigang, et autres
Publié: (2024)
par: Wang, Zhigang, et autres
Publié: (2024)
MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm
par: Zhou, Bowen, et autres
Publié: (2026)
par: Zhou, Bowen, et autres
Publié: (2026)
C-Koordinator: Interference-aware Management for Large-scale and Co-located Microservice Clusters
par: Song, Shengye, et autres
Publié: (2025)
par: Song, Shengye, et autres
Publié: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
par: Tang, Lujie, et autres
Publié: (2024)
par: Tang, Lujie, et autres
Publié: (2024)
Documents similaires
-
Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
par: Mo, Zizhao, et autres
Publié: (2026) -
Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
par: Mo, Zizhao, et autres
Publié: (2024) -
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
par: Wu, Jingfeng, et autres
Publié: (2025) -
PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline Parallelism
par: Kong, Z. Jonny, et autres
Publié: (2025) -
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
par: Hu, Jianmin, et autres
Publié: (2025)