LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Shang, Guo, Junxian, Tang, Haotian, Hu, Qinghao, Xiao, Guangxuan, Tang, Jiaming, Lin, Yujun, Liu, Zhijian, Lu, Yao, Han, Song |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
von: Wang, Yuxin, et al.
Veröffentlicht: (2024)
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
von: Sun, Tingyang, et al.
Veröffentlicht: (2026)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
von: He, Jiaao, et al.
Veröffentlicht: (2024)
von: He, Jiaao, et al.
Veröffentlicht: (2024)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
von: Jayakody, Shakya, et al.
Veröffentlicht: (2026)
von: Jayakody, Shakya, et al.
Veröffentlicht: (2026)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
von: Lawenda, Marcin, et al.
Veröffentlicht: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
von: Lin, Mao, et al.
Veröffentlicht: (2026)
von: Lin, Mao, et al.
Veröffentlicht: (2026)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
Staging Blocked Evaluation over Structured Sparse Matrices
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Efficient Serverless Cold Start: Reducing Library Loading Overhead by Profile-guided Optimization
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2025)
von: Tariq, Syed Salauddin Mohammad, et al.
Veröffentlicht: (2025)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
von: Wahlgren, Jacob, et al.
Veröffentlicht: (2025)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
von: Nicusan, Andrei-Leonard, et al.
Veröffentlicht: (2025)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
von: Debnath, Shimul, et al.
Veröffentlicht: (2026)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
von: Davis, Joshua H., et al.
Veröffentlicht: (2026)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
von: Suo, Jiashun, et al.
Veröffentlicht: (2025)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
von: Battaglini-Fischer, Sándor, et al.
Veröffentlicht: (2025)
von: Battaglini-Fischer, Sándor, et al.
Veröffentlicht: (2025)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
von: Boudaoud, Afif, et al.
Veröffentlicht: (2026)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
von: Liu, Shifang, et al.
Veröffentlicht: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
THEAS: Efficient Power Management in Multi-Core CPUs via Cache-Aware Resource Scheduling
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
von: Muhammad, Said, et al.
Veröffentlicht: (2025)
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
von: Yu, Shan, et al.
Veröffentlicht: (2025)
von: Yu, Shan, et al.
Veröffentlicht: (2025)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
von: Helal, Ahmed E., et al.
Veröffentlicht: (2025)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
FluxSieve: Unifying Streaming and Analytical Data Planes for Scalable Cloud Observability
von: Vogel, Adriano, et al.
Veröffentlicht: (2026)
von: Vogel, Adriano, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
von: Wang, Yuxin, et al.
Veröffentlicht: (2024) -
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
von: Sun, Tingyang, et al.
Veröffentlicht: (2026) -
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
von: Zhou, Qihui, et al.
Veröffentlicht: (2025) -
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025) -
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
von: He, Jiaao, et al.
Veröffentlicht: (2024)