LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Juntao, Wan, Borui, Peng, Yanghua, Lin, Haibin, Wu, Chuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)
LAPS: A Length-Aware-Prefill LLM Serving System
von: She, Jianshu, et al.
Veröffentlicht: (2026)
von: She, Jianshu, et al.
Veröffentlicht: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation
von: Feng, Weiqi, et al.
Veröffentlicht: (2024)
von: Feng, Weiqi, et al.
Veröffentlicht: (2024)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
von: Hankendi, Can, et al.
Veröffentlicht: (2026)
High-Throughput LLM inference on Heterogeneous Clusters
von: Xiong, Yi, et al.
Veröffentlicht: (2025)
von: Xiong, Yi, et al.
Veröffentlicht: (2025)
Laminar: A Scalable Asynchronous RL Post-Training Framework
von: Sheng, Guangming, et al.
Veröffentlicht: (2025)
von: Sheng, Guangming, et al.
Veröffentlicht: (2025)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
von: Cho, Jaehong, et al.
Veröffentlicht: (2026)
von: Cho, Jaehong, et al.
Veröffentlicht: (2026)
LLMServingSim2.0: A Unified Simulator for Heterogeneous Hardware and Serving Techniques in LLM Infrastructure
von: Cho, Jaehong, et al.
Veröffentlicht: (2025)
von: Cho, Jaehong, et al.
Veröffentlicht: (2025)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2026)
HAFLQ: Heterogeneous Adaptive Federated LoRA Fine-tuned LLM with Quantization
von: Su, Yang, et al.
Veröffentlicht: (2024)
von: Su, Yang, et al.
Veröffentlicht: (2024)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
von: Wilkins, Grant, et al.
Veröffentlicht: (2024)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2024)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
von: Jayakody, Shakya, et al.
Veröffentlicht: (2026)
von: Jayakody, Shakya, et al.
Veröffentlicht: (2026)
Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2026)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
von: Argerich, Mauricio Fadel, et al.
Veröffentlicht: (2026)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
von: Huang, Tao, et al.
Veröffentlicht: (2024)
von: Huang, Tao, et al.
Veröffentlicht: (2024)
Block: Balancing Load in LLM Serving with Context, Knowledge and Predictive Scheduling
von: Da, Wei, et al.
Veröffentlicht: (2025)
von: Da, Wei, et al.
Veröffentlicht: (2025)
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
von: Yuan, Yichao, et al.
Veröffentlicht: (2025)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
von: Qiu, Haoran, et al.
Veröffentlicht: (2025)
Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics
von: Zhou, Zijie
Veröffentlicht: (2026)
von: Zhou, Zijie
Veröffentlicht: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
Robust LLM Training Infrastructure at ByteDance
von: Wan, Borui, et al.
Veröffentlicht: (2025)
von: Wan, Borui, et al.
Veröffentlicht: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
Prompt-Aware Scheduling for Low-Latency LLM Serving
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
von: Tao, Yiheng, et al.
Veröffentlicht: (2025)
AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM Serving
von: Xu, Tianhao, et al.
Veröffentlicht: (2026)
von: Xu, Tianhao, et al.
Veröffentlicht: (2026)
Revisiting Parameter Server in LLM Post-Training
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
ProMoE: Fast MoE-based LLM Serving using Proactive Caching
von: Song, Xiaoniu, et al.
Veröffentlicht: (2024)
von: Song, Xiaoniu, et al.
Veröffentlicht: (2024)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
von: Cho, Jaehong, et al.
Veröffentlicht: (2024)
von: Cho, Jaehong, et al.
Veröffentlicht: (2024)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
von: Peng, You, et al.
Veröffentlicht: (2026)
von: Peng, You, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices
von: Zhao, Juntao, et al.
Veröffentlicht: (2024) -
Sandwich: Joint Configuration Search and Hot-Switching for Efficient CPU LLM Serving
von: Zhao, Juntao, et al.
Veröffentlicht: (2025) -
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
von: Zhao, Juntao, et al.
Veröffentlicht: (2025) -
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025) -
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
von: Huang, Shaoyuan, et al.
Veröffentlicht: (2026)