PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Hongbin, Wei, Taosheng, Jiang, Jiazhi, Yan, Hui, Du, Jiangsu, Chen, Zhiguang |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
par: Zhang, Hongbin, et autres
Publié: (2025)
par: Zhang, Hongbin, et autres
Publié: (2025)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
par: Du, Jiangsu, et autres
Publié: (2025)
par: Du, Jiangsu, et autres
Publié: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
par: Wei, Jinhui, et autres
Publié: (2025)
par: Wei, Jinhui, et autres
Publié: (2025)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
par: Bai, Fengyao, et autres
Publié: (2026)
par: Bai, Fengyao, et autres
Publié: (2026)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
par: Guo, Tianyu, et autres
Publié: (2025)
par: Guo, Tianyu, et autres
Publié: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
par: He, Yongchao, et autres
Publié: (2025)
par: He, Yongchao, et autres
Publié: (2025)
Boosting LLM Serving through Spatial-Temporal GPU Resource Sharing
par: Lin, Zejia, et autres
Publié: (2025)
par: Lin, Zejia, et autres
Publié: (2025)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
par: Li, Shengwei, et autres
Publié: (2023)
par: Li, Shengwei, et autres
Publié: (2023)
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
par: Wilkins, Grant, et autres
Publié: (2026)
par: Wilkins, Grant, et autres
Publié: (2026)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
par: Zhao, Alan, et autres
Publié: (2026)
par: Zhao, Alan, et autres
Publié: (2026)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
par: Fan, Jiakun, et autres
Publié: (2025)
par: Fan, Jiakun, et autres
Publié: (2025)
The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency
par: Chen, Huamin, et autres
Publié: (2026)
par: Chen, Huamin, et autres
Publié: (2026)
OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
par: Wu, Siyu, et autres
Publié: (2025)
par: Wu, Siyu, et autres
Publié: (2025)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
par: Wilkins, Grant, et autres
Publié: (2024)
par: Wilkins, Grant, et autres
Publié: (2024)
PolyKAN: Efficient Fused GPU Operators for Polynomial Kolmogorov-Arnold Network Variants
par: Yu, Mingkun, et autres
Publié: (2025)
par: Yu, Mingkun, et autres
Publié: (2025)
Are Bus-Mounted Edge Servers Feasible?
par: Li, Xuezhi, et autres
Publié: (2025)
par: Li, Xuezhi, et autres
Publié: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
par: Han, Yunhe, et autres
Publié: (2026)
par: Han, Yunhe, et autres
Publié: (2026)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
par: Lin, Shouxu, et autres
Publié: (2026)
par: Lin, Shouxu, et autres
Publié: (2026)
KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
par: Zhang, Guilin, et autres
Publié: (2025)
par: Zhang, Guilin, et autres
Publié: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
par: Yu, Minchen, et autres
Publié: (2023)
par: Yu, Minchen, et autres
Publié: (2023)
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
par: Liu, Chongpeng, et autres
Publié: (2025)
par: Liu, Chongpeng, et autres
Publié: (2025)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
par: Zhao, Bohan, et autres
Publié: (2025)
par: Zhao, Bohan, et autres
Publié: (2025)
PICO: Accelerating All k-Core Paradigms on GPU
par: Zhao, Chen, et autres
Publié: (2024)
par: Zhao, Chen, et autres
Publié: (2024)
Performance Characterization of Distributed Deep Learning Strategies: A Quantitative Evaluation of DDP, FSDP, and Parameter Server Architectures on GPU Clusters
par: Ovi, Md Sultanul Islam
Publié: (2025)
par: Ovi, Md Sultanul Islam
Publié: (2025)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
par: Lee, Munkyu, et autres
Publié: (2024)
par: Lee, Munkyu, et autres
Publié: (2024)
Scaling Up Throughput-oriented LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
par: Phung, Thanh Son, et autres
Publié: (2025)
par: Phung, Thanh Son, et autres
Publié: (2025)
Predictable LLM Serving on GPU Clusters
par: Darzi, Erfan, et autres
Publié: (2025)
par: Darzi, Erfan, et autres
Publié: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
par: Wang, Weiye, et autres
Publié: (2026)
par: Wang, Weiye, et autres
Publié: (2026)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
par: Zhang, Haolin, et autres
Publié: (2025)
par: Zhang, Haolin, et autres
Publié: (2025)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
par: Qiao, Yifan, et autres
Publié: (2024)
par: Qiao, Yifan, et autres
Publié: (2024)
Efficiently Executing High-throughput Lightweight LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
par: Phung, Thanh Son, et autres
Publié: (2025)
par: Phung, Thanh Son, et autres
Publié: (2025)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
par: Hu, Cunchen, et autres
Publié: (2024)
par: Hu, Cunchen, et autres
Publié: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
par: Gu, Jianfeng, et autres
Publié: (2025)
par: Gu, Jianfeng, et autres
Publié: (2025)
A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems
par: Wu, Qi, et autres
Publié: (2026)
par: Wu, Qi, et autres
Publié: (2026)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
par: Lin, Mao, et autres
Publié: (2026)
par: Lin, Mao, et autres
Publié: (2026)
LLM-CoOpt: A Co-Design and Optimization Framework for Efficient LLM Inference on Heterogeneous Platforms
par: Kong, Jie, et autres
Publié: (2026)
par: Kong, Jie, et autres
Publié: (2026)
SECO: Secure Inference With Model Splitting Across Multi-Server Hierarchy
par: Chen, Shuangyi, et autres
Publié: (2024)
par: Chen, Shuangyi, et autres
Publié: (2024)
FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
par: Lin, Yanying, et autres
Publié: (2025)
par: Lin, Yanying, et autres
Publié: (2025)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
par: Chung, Euijun, et autres
Publié: (2026)
par: Chung, Euijun, et autres
Publié: (2026)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
par: Da, Wei, et autres
Publié: (2026)
par: Da, Wei, et autres
Publié: (2026)
Documents similaires
-
TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
par: Zhang, Hongbin, et autres
Publié: (2025) -
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
par: Du, Jiangsu, et autres
Publié: (2025) -
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
par: Wei, Jinhui, et autres
Publié: (2025) -
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
par: Bai, Fengyao, et autres
Publié: (2026) -
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
par: Guo, Tianyu, et autres
Publié: (2025)