Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Miao, Xupeng, Oliaro, Gabriele, Zhang, Zhihao, Cheng, Xinhao, Jin, Hongyi, Chen, Tianqi, Jia, Zhihao |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
by: Miao, Xupeng, et al.
Published: (2023)
by: Miao, Xupeng, et al.
Published: (2023)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
by: Sun, Tingyang, et al.
Published: (2026)
by: Sun, Tingyang, et al.
Published: (2026)
A System for Microserving of LLMs
by: Jin, Hongyi, et al.
Published: (2024)
by: Jin, Hongyi, et al.
Published: (2024)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
by: Zhao, Xuanlei, et al.
Published: (2024)
by: Zhao, Xuanlei, et al.
Published: (2024)
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
by: Li, Zikun, et al.
Published: (2025)
by: Li, Zikun, et al.
Published: (2025)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
by: Wang, Yuxin, et al.
Published: (2024)
by: Wang, Yuxin, et al.
Published: (2024)
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
by: Mei, Yixuan, et al.
Published: (2024)
by: Mei, Yixuan, et al.
Published: (2024)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
by: Ather, Hammad, et al.
Published: (2024)
by: Ather, Hammad, et al.
Published: (2024)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
by: Hu, Junhao, et al.
Published: (2024)
by: Hu, Junhao, et al.
Published: (2024)
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
by: Chu, Xiaoyu, et al.
Published: (2025)
by: Chu, Xiaoyu, et al.
Published: (2025)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
by: Shan, Baodi, et al.
Published: (2024)
by: Shan, Baodi, et al.
Published: (2024)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
by: Jayakody, Shakya, et al.
Published: (2026)
by: Jayakody, Shakya, et al.
Published: (2026)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
by: Suo, Jiashun, et al.
Published: (2025)
by: Suo, Jiashun, et al.
Published: (2025)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
by: Scheinert, Dominik, et al.
Published: (2023)
by: Scheinert, Dominik, et al.
Published: (2023)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
by: Xu, Jingwei, et al.
Published: (2025)
by: Xu, Jingwei, et al.
Published: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
by: Lin, Mao, et al.
Published: (2026)
by: Lin, Mao, et al.
Published: (2026)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
by: Suffa, Philipp, et al.
Published: (2024)
by: Suffa, Philipp, et al.
Published: (2024)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
by: Hu, Xiannan, et al.
Published: (2025)
by: Hu, Xiannan, et al.
Published: (2025)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
by: He, Jiaao, et al.
Published: (2024)
by: He, Jiaao, et al.
Published: (2024)
Ridgeline: A 2D Roofline Model for Distributed Systems
by: Checconi, Fabio, et al.
Published: (2022)
by: Checconi, Fabio, et al.
Published: (2022)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
by: Gupta, Ahan, et al.
Published: (2026)
by: Gupta, Ahan, et al.
Published: (2026)
Towards Stream-Based Monitoring for EVM Networks
by: Onica, Emanuel, et al.
Published: (2025)
by: Onica, Emanuel, et al.
Published: (2025)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
by: Nicusan, Andrei-Leonard, et al.
Published: (2025)
by: Nicusan, Andrei-Leonard, et al.
Published: (2025)
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
by: Veeravalli, Bharadwaj
Published: (2026)
by: Veeravalli, Bharadwaj
Published: (2026)
Toward Scalable Docker-Based Emulations of Blockchain Networks for Research and Development
by: Pennino, Diego, et al.
Published: (2024)
by: Pennino, Diego, et al.
Published: (2024)
Towards High-Performance and Portable Molecular Docking on CPUs through Vectorization
by: Accordi, Gianmarco, et al.
Published: (2025)
by: Accordi, Gianmarco, et al.
Published: (2025)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
by: Boudaoud, Afif, et al.
Published: (2026)
by: Boudaoud, Afif, et al.
Published: (2026)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
by: Lawenda, Marcin, et al.
Published: (2025)
by: Lawenda, Marcin, et al.
Published: (2025)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
by: Liu, Shifang, et al.
Published: (2025)
by: Liu, Shifang, et al.
Published: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
by: Wang, Yuxin, et al.
Published: (2023)
by: Wang, Yuxin, et al.
Published: (2023)
An Online Probabilistic Distributed Tracing System
by: Toslali, M., et al.
Published: (2024)
by: Toslali, M., et al.
Published: (2024)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
by: Vatsavai, Sairam Sri, et al.
Published: (2025)
by: Vatsavai, Sairam Sri, et al.
Published: (2025)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
Efficient Serverless Cold Start: Reducing Library Loading Overhead by Profile-guided Optimization
by: Tariq, Syed Salauddin Mohammad, et al.
Published: (2025)
by: Tariq, Syed Salauddin Mohammad, et al.
Published: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
by: Zhuang, Chen, et al.
Published: (2024)
by: Zhuang, Chen, et al.
Published: (2024)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
by: Mathews, Dhanya R, et al.
Published: (2025)
by: Mathews, Dhanya R, et al.
Published: (2025)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
by: Muhammad, Said, et al.
Published: (2025)
by: Muhammad, Said, et al.
Published: (2025)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
by: Proaño, Andrès Rubio, et al.
Published: (2024)
by: Proaño, Andrès Rubio, et al.
Published: (2024)
Similar Items
-
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
by: Miao, Xupeng, et al.
Published: (2023) -
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
by: Sun, Tingyang, et al.
Published: (2026) -
A System for Microserving of LLMs
by: Jin, Hongyi, et al.
Published: (2024) -
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
by: Zhao, Xuanlei, et al.
Published: (2024) -
FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
by: Oliaro, Gabriele, et al.
Published: (2024)